Friday, September 3, 2010

Everything involved with the copy-editing and formatting of an eBook

One of the most important things to remember when editing a manuscript in preparation to publish in eBook format is the text must be as free of errors as possible.  These errors include misspellings, improper punctuation, stray characters from OCR (more about that later), missing words, and other common mistakes.  The enjoyment of a great book can be affected when the text is full of mistakes.
The books I prepare at Crossroad Press come originally in two flavors: original manuscripts in a digital file format such as MS Word and text scanned and OCR'd from the original printed version.  The originals obviously involve less work as there is no cleanup of stray OCR characters, many words with emphasis (underline, bold, or italicized) are already formatted properly, and have possibly already undergone some form of editing in a prior release, unless they are original manuscripts.  At Crossroad Press, we call these original manuscripts "original to digital".

If the original manuscript is not available in file format, then the book has to be chopped up, scanned, and the text read and interpreted by Optical Character Recognition (OCR) software.  Fortunately, OCR software is much better at reading and interpreting text than it was when I first started using it in the mid '90s.  The first step is to remove the binding.  A large blade, like those you'd see at a office supply store or copy/print shop, would be optimal.  After the binding has been removed, the pages are then scanned and OCR'd.  The OCR process then creates, in our case, a series of files in Word format.

That leads to the second step -- reconstruction of the files.  The scanned and interpreted text must be put back in order.  Once they are in order, then the copy-edit process can begin.

Sometimes the OCR software does an excellent job, especially if the original paper is thick and the pages are clean, meaning no smudges, smeared print, water stains, hand written notes, etc. The cleaner files are obviously easier to copy-edit because of fewer OCR'd errors. If the original pages aren't in the best of shape, the OCR engine has more difficulty, which increases the time and effort involved to copy-edit the document. I search for common OCR mistakes such as a capital I being interpreted as a 1, incorrectly displayed ellipses (...) and the like. I have a growing list of common mistakes that I look for in each document. I do the spellchecking when I do the copy-editing, as well as scan for hyphens from words originally hyphenated in the book, such as today being to- (at the end of one line) and day (beginning of next line), and correct those by removing the hyphen. I always do a second scan for those hyphens at the end to see if I've missed any.

The other part of editing is correcting the mistakes that were wrong in the original text. These are normally misspellings or transposed letters. Sometimes the original text might have left off accent marks, like the one over cliché, and I update those as well. Some publishers did a good job with the original edits; however, others seemed like they wanted to rush the book out the door.

The last part of my copy-editing is I take my cleaned up document and compare it to the original text. I ensure that all new lines are indented as they were originally, make sure all quotes are present, apply any special formatting like italics, bold, and underlining. If it's a block of poetry, I will apply a special style to it to try to offset it from the regular text. The final thing is to look for the scene breaks in the original book -- sometimes they have a separator like * * *, other times they have an extra blank line to offset the two paragraphs. To make it consistent, I use a common separator throughout the books.


If the original manuscript is available in digital file format, then the reconstruction process is not needed and neither is the check for common OCR mistakes.  All the other checks (misspellings, hyphen removal, separator replacement, sentence structure, etc.) are applicable to these documents as well.

After all that is done, I send the edited book back to David Wilson at Crossroad Press where he does the final formatting and conversion to the different formats (ePub, PDF, MOBI and PRC). The other step is getting the new cover in place. After the final document formats are ready, they're listed on the Crossroad Press website, then distributed elsewhere.

There's definitely a lot of work that goes into the conversion from start to finish. It's a very rewarding feeling when the book is complete and an author gets to see their work available again, sometimes for the first time in 20+ years.  The process to create and edit a manuscript for conversion to eBook format is not just a simple "save .doc as eBook" task.


If you have any questions about the process or any of the Crossroad titles, let me know and I'll be glad to help.