Skip to main content
FixMyTech

Why Can't I Search or Copy Text in a Scanned PDF?

By

Published

7 min read

Share

Short answer

The page is a picture, not text. A scanner photographs the paper and stores pixels, so there is nothing for search to find. OCR — optical character recognition — reads the image and writes an invisible text layer behind it, after which search and copy work normally. Google Drive does this free on upload.

On this page

Open a scanned contract, press Ctrl+F, search for a word you can plainly see on the page, and get no results. The document is not broken and the search is not failing. There is simply no text in the file to find.

This confuses people because the page clearly shows words. But a scanner does not read — it photographs. What ends up in the PDF is a grid of coloured pixels that happens to look like writing to you and means nothing to the software.

The one-second test

Click and drag across a line of text with your cursor.

  • A highlight appears and follows the words — there is a text layer, and searching should work. If it does not, see the troubleshooting section below.
  • Nothing highlights, or a selection rectangle covers the whole page at once — the page is an image.

That test also tells you what you are dealing with before you try to compress, edit or email the file, because a scanned PDF behaves differently in every one of those tasks. It is the same check that identifies the biggest cause of oversized files in why a PDF is so large.

What OCR actually does

Optical character recognition analyses the image, finds shapes that look like characters, matches them against known letterforms, and uses dictionaries and context to resolve the ambiguous ones. It then writes the result into the PDF as a text layer, invisibly positioned behind the image.

The page looks exactly the same afterwards. What changes is that there is now text sitting underneath each word, aligned to its position, so search finds it and copy retrieves it. Select a word in an OCR’d scan and you are highlighting that invisible layer, which is why the highlight occasionally sits slightly off from the printed word.

Three things decide how well it works:

Resolution. Around 300 DPI is the usual sweet spot for printed text. Below about 200, characters lose the detail that distinguishes similar shapes. Above 400, accuracy stops improving while the file grows.

Contrast and evenness. Black text on white paper, lit evenly. Shadows, a curved page, highlighter pen and coloured backgrounds all cost accuracy.

Alignment. Text should run horizontally. Most OCR engines correct a few degrees of skew; a photo taken at an angle with perspective distortion is much harder.

Getting OCR done without buying anything

Google Drive

The least-effort option and genuinely good.

  1. Upload the PDF to Google Drive.
  2. Right-click → Open with → Google Docs.

Drive runs OCR during the conversion and produces a Google Doc containing the recognised text. Note what you get: an editable document with the text, not your original PDF with a text layer added. For extracting text that is ideal; for keeping the scan’s appearance it is not.

Layout survives poorly. Columns, tables and anything multi-column come out rearranged. For a plain letter or a single-column report it is fine.

Microsoft OneNote

Insert the scan into a OneNote page, right-click the image, and choose the option to copy text from the picture. OneNote also indexes text inside images automatically, so scans stored there become searchable within the notebook without any explicit step.

Good for extracting text from one page; impractical for a 200-page document.

Your phone’s camera

Both mobile platforms recognise text in images at the system level now. Point the camera at a page, or open an existing photo, and selectable text appears on the live image — select, copy, paste.

This is the fastest way to get a paragraph out of a document you are holding. It does not produce a searchable PDF.

Scanner software

Most scanner drivers have an OCR option, often buried under a “searchable PDF” or “document” preset rather than being labelled OCR. Where it exists it is the best route, because the software has the full-resolution scan before any compression has touched it.

If you are scanning with a phone, the capture quality is the thing that decides the outcome, and scanning documents to PDF with your phone covers getting a flat, evenly lit page. Making a scanned PDF searchable goes further into the OCR step itself.

When there is a text layer and search still fails

Occasionally the selection test passes and search still returns nothing.

Ligatures. Some fonts combine fi, fl and similar pairs into a single glyph, and a poorly generated PDF stores that glyph without mapping it back to two characters. Searching for “file” then fails while “le” succeeds.

Missing character mappings. A PDF generated by certain older tools, particularly some LaTeX configurations, can store glyph indices with no indication of which Unicode characters they represent. The page displays correctly and the text is unsearchable and uncopyable, and copy-paste produces gibberish.

Wrong language model during OCR. Accented characters recognised as their unaccented equivalents make searches for the correct spelling fail.

For the first two, the practical workaround is to re-OCR the page as if it were a scan. You lose the original text layer and gain one that is actually mapped to characters.

The accuracy you should expect

Source Realistic accuracy
Clean 300 DPI scan of printed text Very high; occasional errors
Phone photo, flat and evenly lit High on body text, weaker on small print
Faded or photocopied-repeatedly document Variable, needs checking
Fax or low-resolution scan Poor
Handwriting Unreliable, print more than cursive

The characters that go wrong are predictable: 1 and l and I, 0 and O, rn read as m, cl read as d. If you are OCR’ing something where the numbers matter — an invoice, a statement, a reference code — check those by eye rather than trusting the output.

What to expect

For a clean scan of printed text, OCR through Google Drive takes a minute and gives you searchable, copyable text with few errors. Scanner software with a searchable-PDF option does better, because it works from the original scan.

The honest limit is handwriting. Despite steady improvement, handwritten text recognition remains unreliable enough that it cannot be trusted without checking every word, which generally takes longer than typing it. If a scanned document is handwritten, treat it as an image and plan around that rather than hoping a better tool exists.

Frequently asked questions

How do I tell whether a PDF contains real text?
Try to select a word with your cursor. If a highlight appears and follows the line, there is a text layer. If nothing highlights, or a rectangle covers the whole page at once, the page is an image.
Does OCR change how the document looks?
No. The recognised text is written as an invisible layer positioned behind the image, so the page looks identical. You are searching and copying the hidden layer while looking at the original scan.
How accurate is OCR?
On a clean 300 DPI scan of printed text, very accurate — errors are rare enough to be noticeable rather than routine. On a phone photo at an angle, a faded fax or handwriting, accuracy drops sharply, and handwriting in particular is unreliable.
Does adding OCR make the file bigger?
Slightly. The text layer is characters and positions, typically a few kilobytes per page. Some OCR tools also re-compress the page images at the same time, which usually makes the file smaller overall rather than larger.
Can OCR read a document in another language?
Yes, but you generally have to tell it which language. OCR uses a dictionary to resolve ambiguous shapes, so running an English model over a French document produces more errors than running the right one. Most tools have a language setting somewhere in their options.

All PDF guides