Skip to main content
FixMyTech

How to Make a Scanned PDF Searchable

By

Published

7 min read

Share

Short answer

A scan holds a picture of each page, so OCR must add an invisible text layer behind the image. Microsoft Lens on a phone exports a searchable PDF in one step, and OCRmyPDF does it offline. Scan at 200 DPI or higher, compress before you OCR, and verify with Ctrl+F.

On this page

You have a 60-page scanned report, you need the paragraph that mentions a particular supplier, and Ctrl+F returns nothing. You can see the word on the page. The computer cannot.

That is not a bug. A scanned PDF contains no text at all. It contains a photograph of each page, and a photograph of the word “supplier” is, to the computer, indistinguishable from a photograph of a tree.

Why scans are not searchable

When you type a document, each character is stored as a character code plus a reference to a font. The PDF knows that position 47 on page 3 holds the letter “s”. Searching is trivial.

When you scan a document, the scanner records brightness values for several million pixels. There is no “s” anywhere in the file — just a pattern of dark pixels that a human brain recognises as one.

Optical Character Recognition closes the gap. OCR software analyses the pixel patterns, works out which characters they represent, and writes those characters back into the PDF as an invisible text layer positioned exactly behind the corresponding words in the image.

The result is a searchable PDF: it looks exactly like the scan, but Ctrl+F works, you can select and copy text, and screen readers can read it aloud.

Check what you have first

Thirty seconds, and it tells you whether you need OCR at all.

Test Result What it means
Drag the cursor over a word Individual words highlight Real text — already searchable
Drag the cursor over a word Whole page highlights as a block Image only — needs OCR
Drag the cursor over a word Nothing highlights Image only — needs OCR
Ctrl+F a word you can see Found Searchable
Ctrl+F a word you can see Not found Needs OCR, or OCR failed

A document can also be partly searchable — common when scanned pages have been merged into a typed document. Test a few pages, not just the first.

Free ways to run OCR

Google Drive

The most accessible option, available to anyone with a Google account, and it handles a wide range of languages.

  1. Upload the PDF to Google Drive.
  2. Right-click it, choose Open with → Google Docs.
  3. Drive runs OCR and produces a Google Doc containing the recognised text.

The important limitation: this gives you a text document, not a searchable PDF. The layout is approximated and often mangled — columns, tables and headers rarely survive intact. It is excellent for extracting text you need to read, quote or copy; it is not the right tool if you need the original scan to remain visually identical and become searchable.

It also only processes the first ten pages or so of long documents in some cases, so split a long file if you need all of it.

macOS Preview and Shortcuts

Recent versions of macOS have text recognition built into the system, and Preview uses it automatically.

In Preview: open the PDF and try selecting text on a scanned page. On recent macOS versions, Live Text recognises it on the fly and lets you select and copy — without modifying the file. Good for copying a paragraph; it does not make the saved file searchable.

To produce an actual searchable PDF, use a Shortcut:

  1. Open Shortcuts and create a new shortcut.
  2. Add Extract Text from Image or Get Text from Input.
  3. Combine with Make PDF to write the recognised text back.

This takes some assembly. If you want a searchable PDF on macOS without tinkering, a dedicated OCR tool is less effort.

Microsoft Lens

Free on iPhone and Android, and genuinely good. Scan a document and export it as PDF — Lens runs OCR as part of the export and the resulting PDF is searchable. It also exports to Word with a reasonable attempt at preserving the layout.

This is the easiest route if you are scanning the document anyway: scan and OCR in one step rather than scanning now and OCRing later.

OneNote

Also free. Insert the scan into a OneNote page, right-click the image and choose Copy Text from Picture. It gives you the text, not a searchable PDF, but it is quick and already installed on most Windows machines.

LibreOffice Draw

Free and offline, which matters for confidential documents. LibreOffice Draw opens PDFs and lets you edit them, but it does not do OCR on its own — you need Tesseract installed alongside it.

For a genuinely offline, free, proper searchable-PDF result, OCRmyPDF (which wraps Tesseract) is the tool most people end up at. It is a command-line program, so it needs a little comfort with a terminal, but it does exactly the right thing: it adds an invisible text layer behind the original image and leaves the appearance untouched.

Online OCR services

Plenty exist and most work. The caveat is the same one that applies to every online document tool: the documents people need to OCR are overwhelmingly contracts, medical letters, bank statements and official correspondence, and uploading those to an unknown server is a decision worth making deliberately rather than by default.

For anything confidential, use an offline tool.

Adobe Acrobat Pro

Not free, but the benchmark. Scan & OCR → Recognise Text produces the most accurate result, keeps the page image intact, and handles multi-column layouts and tables better than anything else. If you have access to it at work, use it.

Getting good recognition

OCR accuracy depends almost entirely on the input.

Resolution is the main factor. OCR needs at least 200 DPI and prefers 300 DPI on small or faint text. Below 200 DPI the error rate climbs sharply and the errors are confident ones — “rn” read as “m”, “0” as “O”, “1” as “l”.

Straighten the page. Skew of more than a couple of degrees degrades recognition noticeably. Most scanner modes de-skew automatically; photographs taken by hand often do not.

Contrast helps. Clean black text on white paper recognises nearly perfectly. Faded thermal receipts, carbon copies and text over a coloured background or watermark do much worse.

Set the language. Most OCR tools let you specify it, and getting it right substantially improves accuracy on accented characters.

Handwriting is still unreliable. Handwriting recognition has improved but remains far behind printed text on free tools. Do not rely on it for anything that matters.

Compress first, OCR second

The order matters and getting it wrong wastes the work.

Downsampling re-renders the pages as new images, which discards any existing text layer. If you OCR a scan and then compress it, the text layer is thrown away and you are back where you started.

So: compress first with our PDF compressor, then run OCR on the compressed file. The text layer is only a few kilobytes per page, so it adds effectively nothing to the final size.

One adjustment: compress to 200 DPI rather than 150 if the document needs to be searchable. 150 DPI is fine for reading but marginal for OCR. The guidance on where readability and recognition thresholds sit is in compress a scanned PDF without making it unreadable.

Verify it worked

Do not assume. OCR fails quietly, and a PDF with a bad text layer looks identical to one with a good one.

1. The Ctrl+F test. Open the PDF, press Ctrl+F (Cmd+F on Mac) and search for a distinctive word you can see on page 1. Then try a word from the middle of the document and one from the last page — some tools only process the first few pages.

2. The selection test. Drag your cursor across a line of text. Individual words should highlight, roughly aligned with the printed words. If the highlight boxes are wildly offset from the text, the layer exists but is misaligned, which breaks copy-paste.

3. The copy test. Select a paragraph, copy it, and paste it into a text editor. This is where recognition errors become visible — you will see the confusions OCR made. If you need a word count on what came out, our word counter handles it.

4. Check the numbers. Recognition errors in digits are the most damaging and the least obvious. Spot-check any reference numbers, dates or amounts against what is printed on the page.

What not to do

  • Do not OCR before compressing. The compression step discards the text layer and you repeat the work.
  • Do not assume OCR output is accurate enough to act on. For anything legal, financial or medical, read the original image rather than trusting the recognised text. OCR errors look like correct text.
  • Do not OCR below 200 DPI and expect usable results. Go back to the original scan if the compressed copy is too low.
  • Do not upload confidential scans to free online OCR services. This is the single category of document most likely to be sensitive and most likely to be uploaded without thinking.
  • Do not replace the original with a reflowed Word or Docs output unless you have checked it. Layout mangling is common and usually irreversible once the original is gone.

Frequently asked questions

Why can I not search a scanned PDF?
A scanned PDF contains no text at all — just brightness values for several million pixels. Typed documents store each character as a character code plus a font reference, which is why searching works. OCR closes the gap by recognising the pixel patterns and writing an invisible text layer behind the image.
What is the easiest free way to make a scan searchable?
Microsoft Lens, free on iPhone and Android, runs OCR as part of its PDF export, so the resulting file is searchable. For an offline result on a computer, OCRmyPDF (which wraps Tesseract) adds an invisible text layer behind the original image and leaves the appearance untouched.
Does Google Drive produce a searchable PDF?
No. Uploading the PDF and opening it with Google Docs runs OCR, but the output is a text document rather than a searchable PDF, and the layout is approximated — columns, tables and headers rarely survive. It is excellent for extracting text to read or quote, not for keeping the scan visually identical.
Should I compress before or after running OCR?
Compress first, then OCR. Downsampling re-renders the pages as new images and discards any existing text layer, so OCRing first wastes the work. The text layer is only a few kilobytes per page, so adding it afterwards costs essentially nothing in file size.
What resolution does OCR need?
At least 200 DPI, and 300 DPI on small or faint text. Below 200 DPI the error rate climbs sharply and the mistakes are confident ones — "rn" read as "m", "0" as "O", "1" as "l". If you are compressing a document you intend to OCR, use 200 DPI rather than 150.

All PDF guides