How to Make a Scanned PDF Searchable
Short answer
A scan holds a picture of each page, so OCR must add an invisible text layer behind the image. Microsoft Lens on a phone exports a searchable PDF in one step, and OCRmyPDF does it offline. Scan at 200 DPI or higher, compress before you OCR, and verify with Ctrl+F.
On this page
You have a 60-page scanned report, you need the paragraph that mentions a particular supplier, and Ctrl+F returns nothing. You can see the word on the page. The computer cannot.
That is not a bug. A scanned PDF contains no text at all. It contains a photograph of each page, and a photograph of the word “supplier” is, to the computer, indistinguishable from a photograph of a tree.
Why scans are not searchable
When you type a document, each character is stored as a character code plus a reference to a font. The PDF knows that position 47 on page 3 holds the letter “s”. Searching is trivial.
When you scan a document, the scanner records brightness values for several million pixels. There is no “s” anywhere in the file — just a pattern of dark pixels that a human brain recognises as one.
Optical Character Recognition closes the gap. OCR software analyses the pixel patterns, works out which characters they represent, and writes those characters back into the PDF as an invisible text layer positioned exactly behind the corresponding words in the image.
The result is a searchable PDF: it looks exactly like the scan, but Ctrl+F works, you can select and copy text, and screen readers can read it aloud.
Check what you have first
Thirty seconds, and it tells you whether you need OCR at all.
| Test | Result | What it means |
|---|---|---|
| Drag the cursor over a word | Individual words highlight | Real text — already searchable |
| Drag the cursor over a word | Whole page highlights as a block | Image only — needs OCR |
| Drag the cursor over a word | Nothing highlights | Image only — needs OCR |
| Ctrl+F a word you can see | Found | Searchable |
| Ctrl+F a word you can see | Not found | Needs OCR, or OCR failed |
A document can also be partly searchable — common when scanned pages have been merged into a typed document. Test a few pages, not just the first.
Free ways to run OCR
Google Drive
The most accessible option, available to anyone with a Google account, and it handles a wide range of languages.
- Upload the PDF to Google Drive.
- Right-click it, choose Open with → Google Docs.
- Drive runs OCR and produces a Google Doc containing the recognised text.
The important limitation: this gives you a text document, not a searchable PDF. The layout is approximated and often mangled — columns, tables and headers rarely survive intact. It is excellent for extracting text you need to read, quote or copy; it is not the right tool if you need the original scan to remain visually identical and become searchable.
It also only processes the first ten pages or so of long documents in some cases, so split a long file if you need all of it.
macOS Preview and Shortcuts
Recent versions of macOS have text recognition built into the system, and Preview uses it automatically.
In Preview: open the PDF and try selecting text on a scanned page. On recent macOS versions, Live Text recognises it on the fly and lets you select and copy — without modifying the file. Good for copying a paragraph; it does not make the saved file searchable.
To produce an actual searchable PDF, use a Shortcut:
- Open Shortcuts and create a new shortcut.
- Add Extract Text from Image or Get Text from Input.
- Combine with Make PDF to write the recognised text back.
This takes some assembly. If you want a searchable PDF on macOS without tinkering, a dedicated OCR tool is less effort.
Microsoft Lens
Free on iPhone and Android, and genuinely good. Scan a document and export it as PDF — Lens runs OCR as part of the export and the resulting PDF is searchable. It also exports to Word with a reasonable attempt at preserving the layout.
This is the easiest route if you are scanning the document anyway: scan and OCR in one step rather than scanning now and OCRing later.
OneNote
Also free. Insert the scan into a OneNote page, right-click the image and choose Copy Text from Picture. It gives you the text, not a searchable PDF, but it is quick and already installed on most Windows machines.
LibreOffice Draw
Free and offline, which matters for confidential documents. LibreOffice Draw opens PDFs and lets you edit them, but it does not do OCR on its own — you need Tesseract installed alongside it.
For a genuinely offline, free, proper searchable-PDF result, OCRmyPDF (which wraps Tesseract) is the tool most people end up at. It is a command-line program, so it needs a little comfort with a terminal, but it does exactly the right thing: it adds an invisible text layer behind the original image and leaves the appearance untouched.
Online OCR services
Plenty exist and most work. The caveat is the same one that applies to every online document tool: the documents people need to OCR are overwhelmingly contracts, medical letters, bank statements and official correspondence, and uploading those to an unknown server is a decision worth making deliberately rather than by default.
For anything confidential, use an offline tool.
Adobe Acrobat Pro
Not free, but the benchmark. Scan & OCR → Recognise Text produces the most accurate result, keeps the page image intact, and handles multi-column layouts and tables better than anything else. If you have access to it at work, use it.
Getting good recognition
OCR accuracy depends almost entirely on the input.
Resolution is the main factor. OCR needs at least 200 DPI and prefers 300 DPI on small or faint text. Below 200 DPI the error rate climbs sharply and the errors are confident ones — “rn” read as “m”, “0” as “O”, “1” as “l”.
Straighten the page. Skew of more than a couple of degrees degrades recognition noticeably. Most scanner modes de-skew automatically; photographs taken by hand often do not.
Contrast helps. Clean black text on white paper recognises nearly perfectly. Faded thermal receipts, carbon copies and text over a coloured background or watermark do much worse.
Set the language. Most OCR tools let you specify it, and getting it right substantially improves accuracy on accented characters.
Handwriting is still unreliable. Handwriting recognition has improved but remains far behind printed text on free tools. Do not rely on it for anything that matters.
Compress first, OCR second
The order matters and getting it wrong wastes the work.
Downsampling re-renders the pages as new images, which discards any existing text layer. If you OCR a scan and then compress it, the text layer is thrown away and you are back where you started.
So: compress first with our PDF compressor, then run OCR on the compressed file. The text layer is only a few kilobytes per page, so it adds effectively nothing to the final size.
One adjustment: compress to 200 DPI rather than 150 if the document needs to be searchable. 150 DPI is fine for reading but marginal for OCR. The guidance on where readability and recognition thresholds sit is in compress a scanned PDF without making it unreadable.
Verify it worked
Do not assume. OCR fails quietly, and a PDF with a bad text layer looks identical to one with a good one.
1. The Ctrl+F test. Open the PDF, press Ctrl+F (Cmd+F on Mac) and search for a distinctive word you can see on page 1. Then try a word from the middle of the document and one from the last page — some tools only process the first few pages.
2. The selection test. Drag your cursor across a line of text. Individual words should highlight, roughly aligned with the printed words. If the highlight boxes are wildly offset from the text, the layer exists but is misaligned, which breaks copy-paste.
3. The copy test. Select a paragraph, copy it, and paste it into a text editor. This is where recognition errors become visible — you will see the confusions OCR made. If you need a word count on what came out, our word counter handles it.
4. Check the numbers. Recognition errors in digits are the most damaging and the least obvious. Spot-check any reference numbers, dates or amounts against what is printed on the page.
What not to do
- Do not OCR before compressing. The compression step discards the text layer and you repeat the work.
- Do not assume OCR output is accurate enough to act on. For anything legal, financial or medical, read the original image rather than trusting the recognised text. OCR errors look like correct text.
- Do not OCR below 200 DPI and expect usable results. Go back to the original scan if the compressed copy is too low.
- Do not upload confidential scans to free online OCR services. This is the single category of document most likely to be sensitive and most likely to be uploaded without thinking.
- Do not replace the original with a reflowed Word or Docs output unless you have checked it. Layout mangling is common and usually irreversible once the original is gone.