OCR

Why can't I search a scanned PDF? (and how to fix it)

Because a scan is a photograph. Your device sees an image of a page, not letters, so there is nothing for search to match. The fix is OCR — optical character recognition — which reads the shapes in the image and writes an invisible text layer underneath. The page looks exactly the same afterwards, but it becomes searchable, selectable and copyable.

What OCR actually produces

A searchable scan is two layers: the original image on top, and machine-read text hidden behind it. That is why the page still looks like a scan while search works. It also means OCR never changes how the document looks — if a tool alters the appearance, it did something else.

What determines accuracy

Three things, in order of impact: resolution (300 dpi or better), straightness (a skewed page confuses letter boundaries), and lighting (shadows across the page and glare from a phone camera both hurt).

Photograph a page flat, from directly above, with even light and no shadow from your own hand. That single habit does more for accuracy than switching tools.

What still fails

Handwriting is unreliable in every OCR engine — cursive especially. Very small print, heavily stylised fonts, tables with no visible borders, and text over photographs all reduce accuracy. Multi-column layouts sometimes come out with the columns interleaved.

Treat OCR output as searchable, not as authoritative. For anything where a wrong digit matters — an invoice total, a dosage, an account number — read the original.

Privacy note worth knowing

Many OCR services upload your document to a server to process it. For a bank statement or a medical record that is a meaningful decision, not a technical detail. On-device OCR avoids the question entirely.

PDF Studio – Reader & Tools

Our iOS app for exactly this. Runs on your device — no account, no tracking.

See the app →