OCR PDF
Make a scanned PDF searchable by adding a text layer behind the image. Processed on the server; your file is not stored.
Processed on our server and deleted immediately — your file is never stored.
Pages that already have text are left untouched.
About OCR PDF
OCR PDF reads the words in a scanned document and adds them back as a real, invisible text layer underneath the original image. The page looks exactly the same, but you can now search it, select and copy from it, and let other software extract its content. Free, and your file is deleted straight after processing.
A scanned PDF is a photograph of paper. To you it obviously says something; to your computer it is a grid of coloured dots with no words in it at all. That is why pressing Ctrl+F in a scanned contract finds nothing, why you cannot copy a paragraph out of it, and why converting it to Word produces an empty document. Optical character recognition is what closes that gap.
The practical payoff shows up in archives. A folder of two hundred scanned invoices is nearly useless until it is searchable, at which point finding the one from a specific supplier takes seconds instead of an afternoon. Lawyers run OCR over discovery bundles so a name can be traced across thousands of pages. Researchers do it to old journal articles. And anyone who needs a scanned document to comply with accessibility requirements needs OCR, because a screen reader cannot read an image of text to a blind user.
How to OCR PDF
- Upload the scanned PDF you want to make searchable.
- Wait while each page is analysed and the characters are recognised — this is slower than most tools, and long documents legitimately take a while.
- Download the new PDF, which looks identical to the original.
- Open it and press Ctrl+F, then search for a word you can see on the page.
- If the text is found and highlighted, the OCR layer is working.
Tips & common problems
Scan quality decides accuracy
300 DPI is the sweet spot for OCR. Below about 200 DPI, characters blur into each other and error rates climb sharply. If results are poor, rescanning at a higher resolution helps far more than reprocessing the same file.
Straighten crooked pages first
Text recognition assumes roughly horizontal lines. A page scanned at an angle produces noticeably worse results, so fix the rotation with Organize PDF before running OCR.
Handwriting will not work
This recognises printed type. Handwritten notes, signatures and calligraphy are not reliably readable, and you should expect nonsense rather than a partial result.
OCR then convert, in that order
If your goal is an editable Word file from a scan, run OCR first and feed the searchable output into PDF to Word. Converting a raw scan directly produces an empty document.
Frequently asked questions
What does OCR do to a PDF?
It recognises the printed characters in the page images and stores them as an invisible text layer aligned behind the picture. The page looks unchanged, but the words are now real text that can be searched, selected and copied.
Which languages are supported?
Standard Latin-script languages such as English, French, German, Spanish, Italian and Portuguese work out of the box. Additional language packs including non-Latin scripts can be enabled on the server if you need them.
Will my scanned pages look different afterwards?
No. The original page image is kept exactly as it was and the text is added underneath it, so the document looks identical — it simply gains abilities it did not have before.
How accurate is the text recognition?
On a clean 300 DPI scan of ordinary printed text, accuracy is typically well above 98%. Faint photocopies, unusual fonts, tight columns and skewed pages all reduce it.
Why is OCR slower than other tools?
Every page has to be rendered, analysed and character-matched, which is far more work than copying pages around. A long document can take a minute or more, and that is normal rather than a sign something has stalled.
Is my scanned document kept on your server?
No. It is processed and then deleted immediately, with nothing logged or retained. Since scans are often contracts and identity documents, this matters — the file exists on our server only for the seconds it takes to process.