Docly practical guide

How to OCR a Scanned PDF

A PDF from a scanner usually contains nothing but page images: you cannot search it and you cannot copy from it. OCR recognises the characters and adds an invisible text layer behind the page, which is what turns a scan into a usable document.

6 minute readUpdated 26 July 2026Docly

Steps to OCR a PDF

Add the file, select the document's main language and start processing. If the document mixes two languages, the combined option is safer; for single-language documents, choosing that language improves accuracy.

The quickest way to confirm the result worked is to search a PDF reader for a word you can see on the page. If it is found, the text layer exists. The page itself looks unchanged, because the recognised text sits invisibly behind the image.

What determines OCR accuracy

Scan quality matters more than the software. A page that is straight, captured at sufficient resolution and lit evenly without shadows will always outperform a better tool fed a poor image. No setting rescues a very low-resolution scan.

Choosing the right document language has a direct effect, since language-specific characters are the first to be misread under the wrong model. For documents photographed with a phone, correcting the perspective and cleaning the page before OCR gives noticeably better results than running recognition on the raw photo.

Review the output

OCR is a convenience, not a certification. Compare names, dates, amounts, account numbers and identification numbers against the page image. The usual confusions are predictable: zero and capital O, one and lowercase l, five and S, eight and B.

For documents full of figures, recalculating a total and comparing it with the printed total is far faster than reading cell by cell. For low-quality scans there is no setting that replaces human review.

What OCR unlocks

Once a text layer exists, the document stops being just searchable and becomes usable as input for other tools. Converting to Word or Excel, translating into another language, comparing two versions and redacting text all require a text layer to work at all.

For scans destined for an archive, running OCR and then converting to PDF/A leaves a copy that still opens years later and can still be searched, which is considerably more valuable than storing page images alone.

Quick checklist

  • Capture pages straight, evenly lit and at sufficient resolution.
  • Choose the correct document language.
  • Manually verify important numbers and names.
  • For archives, run OCR first and then convert to PDF/A.

Tools in this guide

Frequently asked questions

What does OCR do?

It turns writing inside an image into searchable and selectable text.

Can OCR read handwriting?

OCR is more reliable for printed text; handwriting results depend on the script and image quality.

Related guides

View all guides