OCR PDF
Recognise text in scanned pages so the PDF becomes searchable and selectable.
OCR PDF online, free and private
Scanned documents and photographed pages look like text but are really just images — you can't select a sentence, copy a reference number, or search for a name. OCR (optical character recognition) fixes that by reading the shapes on each page and recognising the letters and words behind them. This tool runs OCR over your PDF and adds an invisible text layer on top of the original scan, turning a flat picture into a document you can search, select, and copy from while it still looks exactly the same.
Everything happens inside your browser. Your PDF is rendered and analysed on your own device, and only the small language model needed for recognition is fetched from a CDN the first time you use it. The document itself is never uploaded, so contracts, medical records, and archived paperwork stay entirely on your machine.
How to ocr pdf in your browser
- 1
Add your scanned PDF
Click ‘Select PDF file’ or drop it onto the page. Image-only scans and photographed documents are exactly what this tool is built for.
- 2
Choose the document language
Pick the language that matches the text on the page so recognition is as accurate as possible. The first time you use a language, a small model downloads from a CDN; after that it works without re-fetching.
- 3
Run the OCR
Press the OCR button and Tesseract processes each page in turn, reading the characters and building a hidden text layer. Progress is shown as it moves through the document.
- 4
Download the searchable PDF
When it finishes, download your new PDF. It looks identical to the original, but you can now select, copy, and search the text — and so can search engines and PDF readers.
Why use PDFToolz for this
Your documents stay private
Recognition runs on-device with Tesseract compiled to WebAssembly. The only thing that touches the network is the language model, downloaded from a CDN — your actual pages never leave the browser, which matters for sensitive scans like IDs, statements, and legal files.
The original scan is preserved
We don't replace your page with recognised text. The invisible text layer sits behind the original image, so the document still looks exactly as it was scanned while becoming fully selectable and searchable underneath.
Multiple languages supported
Choose the language of your document and the matching Tesseract model handles the alphabet and word patterns for you. Recognition accuracy improves noticeably when the selected language matches the text on the page.
Free with no limits
OCR is completely free — no account, no watermark, and no page cap. Because the heavy lifting happens on your own hardware, there are no per-page charges to pass on to you.