LoudReader Logo

How to Listen to Scanned PDF Books

Built by the developer of LoudReader

Last updated:

Selectable text means the PDF will read. Image-only means it will not. Test before you import.

Why do some scanned PDFs work with TTS and others don't?

A scanned PDF is a container. What is inside the container determines whether TTS works. There are two possibilities:

  • PDF with a text layer. Behind the page image, invisible to your eyes, is a layer of selectable text that OCR software generated when the scan was created. This text is what TTS engines read. You can test for it by trying to select a word with your cursor. If it highlights and you can copy it, the text layer exists.
  • Image-only PDF. The PDF contains page images and nothing else. Each page is a photograph of printed text, not text itself. Clicking on a word does nothing, or selects a rectangular region of the image. No text means nothing for TTS to read.

This distinction is the whole article. Everything downstream depends on which type of PDF you have.

How do I test if my PDF can be read aloud?

The test takes five seconds and works on any device:

  1. Open the PDF in Preview (Mac), the Files app (iPhone), or any PDF viewer.
  2. Try to select a word by clicking and dragging over it with your cursor or finger.
  3. If the word highlights, the PDF has a text layer. It will work in LoudReader.
  4. If nothing highlights and you get a crosshair or a rectangular selection, the PDF is image-only. It will not produce audio in LoudReader.

This is the single most useful thing you can do before importing a scanned PDF into any TTS app, not just LoudReader. Every TTS engine needs actual text to work with.

What do I do with an image-only scanned PDF?

You need to run OCR on it. The goal is to add a text layer to the existing page images so that a TTS engine has something to read. Here are the options, from simplest to most thorough:

  • Apple Live Text (free, built in). Open the image-only PDF in Preview on Mac. Click and drag to select the area with text. If the text is clear enough, Live Text recognizes it and lets you copy it. Paste the copied text into a text editor, save as a new PDF, and import into LoudReader. This works for short documents but is tedious for a 300-page book.
  • Adobe Acrobat Pro (paid). Open the PDF, go to Scan and OCR, then Recognize Text. Acrobat processes the entire document and adds a text layer behind each page image. The result works in any TTS app, including LoudReader.
  • Tesseract OCR (free, open source). Tesseract is a command-line OCR engine. It requires some setup but produces good results for clear, well-lit scans. The command is something like tesseract scan.pdf output pdf. The output is a new PDF with a text layer.

After OCR, test the new PDF by selecting text. If text is selectable, import into LoudReader and press play. For more on the import workflow, see listening to PDFs on iPhone.

How good will the reading sound after OCR?

The voice quality is the same natural neural voice LoudReader always uses. But OCR accuracy varies a lot, and the errors become audible:

  • Clean scans produce clean audio. A well-lit, high-resolution scan of a modern printed book, run through good OCR software, might have 99% accuracy. One error per hundred words is barely noticeable when listening.
  • Poor scans produce noisy audio. A faded, skewed, or low-resolution scan of an old paperback with tight margins produces OCR errors that the TTS voice reads as garbled text. "cl" becomes "d", "rn" becomes "m", and sentences containing these errors sound like a book full of typos. This is not a LoudReader limitation. It is an OCR limitation, and no TTS engine can fix bad input text.

The honest advice: if you are scanning books for TTS listening, invest in a good scan first. Flatbed scanner, high resolution, good lighting, and OCR software with a preview step so you can correct common errors. A clean scan takes more time upfront but makes the difference between a listenable audiobook and a frustrating one.

Where do I find PDFs that already have text layers?

Most professionally produced PDFs from the last two decades include a text layer because it enables search, copy, and accessibility features. These sources reliably produce text-layer PDFs:

  • The Internet Archive. Millions of scanned public domain books with OCR text layers already applied. Download the PDF and test text selection. Most work with LoudReader.
  • Google Books. Public domain books on Google Books are available as PDFs with text layers. The OCR quality is generally good.
  • Academic databases. JSTOR, PubMed, and university repositories distribute PDFs with text layers. These are designed for text extraction and work reliably with TTS.
  • Library scans. Many public and university libraries scan their collections with OCR enabled. The resulting PDFs are searchable and listenable.

For books that are already in the public domain, turning any book into an audiobook covers the full path from file to audio. The same principles apply to scanned PDFs with text layers.

Frequently asked questions

Can LoudReader read scanned PDFs?

It depends entirely on the PDF. If the scan includes a text layer (selectable, highlightable text behind the image), yes. LoudReader reads the text layer and ignores the image overlay. If the scan is image-only with no selectable text, no. LoudReader has no OCR capability and cannot extract text from images. The PDF will appear in the app but produce no audio because there is nothing to read.

How do I check if my scanned PDF has a text layer?

Open the PDF in any PDF viewer (Preview on Mac, the Files app on iPhone, or Adobe Reader). Try to select a word by clicking and dragging over it. If the text highlights and you can copy it, the PDF has a text layer. If you click and nothing selects, or you get a crosshair cursor that selects a rectangular region of the image, the PDF is image-only. Only the first type works with LoudReader.

What can I do with an image-only scanned PDF?

You need to run OCR (optical character recognition) on it first. Apple's built-in Live Text can extract text from images in Preview, Photos, and Quick Look. Open the scanned PDF in Preview, select the area with text, and copy it. Paste into a text editor and export as a new PDF. For book-length scans, a dedicated OCR tool like Adobe Acrobat Pro or the open-source Tesseract engine produces better results with page structure preserved.

Can LoudReader read old books I scanned myself?

If you scanned them with software that includes OCR (most scanning apps and all-in-one printers offer this as a checkbox), the resulting PDF likely has a text layer. Test by selecting text in the PDF. If text is selectable, LoudReader reads it aloud. If you scanned as image-only (JPEG or TIFF embedded in a PDF container), the test will fail and you need to OCR the file first.

How good is the TTS quality on scanned PDFs?

The voice quality is the same natural neural voice LoudReader always uses. But OCR accuracy affects what words the TTS engine sees. OCR errors (sc as so, rn as m, cl as d) produce misread words. A cleanly OCR'd book reads well. A messy scan with lots of OCR errors sounds like a book full of typos. There is no way around this: the output is only as good as the OCR that produced the text layer.

Where do I find PDFs that already have a text layer?

Publications from academic databases (JSTOR, PubMed), Google Books, the Internet Archive, and most modern ebook platforms include a text layer by default. Library scans of public domain books almost always have a text layer because adding OCR is the difference between a useful archive and a stack of unsearchable images. If you are downloading a PDF to listen to, prefer these sources. If you are scanning your own books, enable OCR in your scanning software before saving.

Hear your scanned books read aloud

Test for selectable text, import into LoudReader, and press play. Natural voices, offline, no account.

Download on theApp Store

Free download for Mac and iPhone · works on iPad too

Keep reading

Still have questions? Get in touch