Large corporations have been buying old books. They cut the bindings and covers off, and then feed the stacked pages into high-speed scanners to digitize the books. At the end of the process, the old books have been destroyed. This would not be an enormous tragedy if the scans were used to reprint the old books, and make their digital images available to the public. However, the scans are being used to train AI large language models. Whether from indifference or to keep competing AI products from using the same data, the scans are not being made available for use by anyone. Many of the books being scanned are rare, and since no record has been kept of what books are being scanned, we do not know what has been lost thus far.
Fortunately, it is possible to digitize old books at home without destroying them. At the end of the process, you can make them available on websites like The Internet Archive; Reprint them on archival paper with an inkjet printer; Email them to others who are interested in them. You probably have most of the equipment you need– A home computer, and a printer with a flatbed scanner.
You may even have the software you need– A program that runs your scanner, such as NAPS2, AABBY, Adobe, or Xsane; Picture processing software like Microsoft Office Picture Manager, ImageMagick, RawTherapee, or Adobe Photoshop; PDF Editing Software like Qoppa PDF Studio Pro, Adobe Acrobat, or FoxIt. Prices range from free to monthly payment forever, and software of each type can be found for Windows, Linux, and Mac.
You don’t need a lot of features to digitize a book, so top-of-the-line software is not necessary. The method I use is based on the software and operating systems I already had. Your resources may be different or (hopefully!) better than mine. What I describe may allow you to figure out how to do this with what you already own.
Books that are in the public domain (generally more than 80 years from their creation date) are often in rough shape. If they were printed on cheap acid paper, the pages become brown and brittle with increasing age. This makes them easy to damage and increasingly difficult to read. With careful handling, you can reverse most of the effects of aging with image processing tools while you avoid damaging the book.
Before You Start
See if someone else has already scanned your book. Check The Internet Archive. Google your book title, and see if a PDF or an ePub comes up.
If you find an awesome scan that you can download, great! Pick another old book, see if you can find it online. If not, scan THAT one!
Even if you can find a scanned copy– Is it a high quality scan? Are all the pages there? Could your re-scan be better? Use your judgement.
General Process
We want to scan the pages to PNG picture files.
We want to improve the images while they are still picture files.
We convert them into PDF images, and link them together in the same order as the pages in the book. It’s as easy as that!
We don’t want to do each page individually. It’s better to process all the pages through each step, since each step may need a different software package, and perhaps even a different operating system.
Step 1. Scan pages to PNG picture files
I use Xsane for scanning, a free app that works under Linux.
Create a directory to receive the scanned images.
Open Xsane.
Set the directory and filename. Example; Documents/Book01/BookPage-000.png
Xsane will assign this filename to the first image. It will increment the 3-digit
number at the end of the file name, and will put all the images in the same directory.
The next image will be named ‘BookPage-001.’ The one after that will be named
‘BookPage-002' and so on.
Set resolution to 300 DPI (Dots Per Inch).
Make it a Colour scan (not black and white)
Clean the Flatbed Scanner glass
With glass cleaner and a microfibre cloth if possible.
Be sure the glass is dry before putting the book on it.
Put the book face down on the Scanner glass.
Use firm pressure to flatten the pages, but be careful not to break the book spine.
You will scan two facing pages at a time.
Do a preview scan for the first time only.
Drag the dotted line that defines the scan area to match the area of the two pages.
Try to go to the edge of the pages, but not beyond them.
When the preview has defined the scan area, it will scan the same area for all the rest of the pages (unless you re-set it with preview).
Scan each set of two pages, twice. This makes a future step easier.
The page just scanned appears for confirmation.
Save the two-page set. Scan it again, save it again. Flip to the next page.
Step 2. Improve the Images
You now have a directory with perhaps several hundred images, each displaying two
facing pages. This took a lot of effort!
Make a second directory, and copy all the dual pages into the second directory.
If something terrible happens to those files, you don’t want to scan them again!
This second directory could be on a flash drive.
Close Linux, and start Windows. Plug in the flash drive.
Start Microsoft Windows Picture Manager
Select the second directory from within Picture Manager.
Picture Manager will only look at the images in this second directory, not all of your
pictures everywhere.
Select [Edit] tab in upper right.
Click down arrow, Select [Color] tab.
Go to Saturation slider at the bottom.
Slide it all the way to the left to remove all colour.
This changes the page colour from brown to white, or light grey.
[Ctrl] S to save the image.
Use right arrow (bottom, centre) to go to the next page. Repeat.
Close Microsoft Picture Manager. Save any files that may not have been saved yet.
Step 3. Crop to One Page Image per File
You now have a directory with new-looking page images. It took a lot of effort!
Make a THIRD directory, and copy all the files from directory 2 into directory 3.
Open Microsoft Picture Manager again.
Set the directory to Directory 3.
Select [Edit], scroll down to [Crop].
Remember that there are two (2) copies of each pair of facing pages.
Drag the Crop rectangle around the page on the left, and [Save].
Arrow to the next image.
Drag the Crop rectangle around the page on the right, and [Save].
Work your way through all the files. At the end, you will have PNG files that each contain
a single page. The pages naturally fall into order.
Did you delete the wrong page or crop them backwards? Don’t worry. You can copy a
file from the Second Directory and re-do it, because you DID do step 3 to copy those files
into the 3rd directory, right? Good!
Close Microsoft Picture Manager and Save any straggler files.
Step 4. Convert to PDF images and Merge into a Multi-Page PDF Book
The 4th directory is filled with PNG images that each contain only one page of the book.
Really, it took a lot of effort to get them into this form!
Open Qoppa PDF Studio Pro.
Select ‘Create PDF from Multiple Files.’
Find the 4th directory, and Select ‘Add Files.’
Be sure all the file names are in order.
If needed, you can use the blue arrows at the bottom of the list to move files around.
Click [Start] at the bottom.
Qoppa PDF Studio Pro converts the PNG files to PDFs and chains them together.
[Save] to save your scanned book!
OPTION: OCR
You can run Optical Character Recognition (OCR) to add an underlying layer to the
PDF book that will allow you to highlight text, and export it to a document file, like
Microsoft Word. OCR tends not to be 100% accurate, but with careful proofreading you could find and edit the errors in the text, and re-create the manuscript in a word processor without having to re-type the whole book.
Step 5. Share Your Old Book
Congratulations! You have preserved an important historical book for future generations! Now that you have it, share it with others–
Upload it to The Internet Archive.
At this writing, that’s https://archive.org/details/books .
Add a concise description and keywords to make it easy for people to find.
See if your local library accepts PDF book uploads.
Upload it to the US Library of Congress
https://www.loc.gov/programs/cataloging-in-publication/ebooks-program/ebooks-application-process/
Library of Congress prefers the ePub format, but will accept PDF files. Your PDF software may be able to export your book to ePub format.
Upload it to Library and Archives Canada https://www.canada.ca/en/library-archives/services/publishers/legal-deposit/digital-publications.html#a1
With registration, it may be possible to donate your scanned books to Library and Archives Canada, the Canadian National Library.
If you specify ‘Open Access’ for your upload, anyone may view and download your donation through the LAC website; https://www.canada.ca/en/library-archives/collection/search.html
Upload to your own National Library
Most countries have them!
Print your book.
Try to find Archival-quality paper, and print it on an inkjet printer to create a lasting copy of your book. Laser printer ink can flake off of the page in 10 to 20 years. There may be a more durable printing method for archival documents. Do the research.
If possible, print the book on a single side of each page, and have it bound.
Properly printed and stored, a paper book can last for many centuries!
That’s the outline, and as far as I have gotten with it.
The App ‘Image Magick’ is a command-line image-processing program available for Linux and Windows that may be able to do a lot of the cropping and contrast operations on the entire directory of images at one time! If you do this a lot, it may be worth trying. Best of luck to all my fellow digitizers!
Comments
Post a Comment