PDF OCR & Recognition — Scan to Searchable PDF, Text & Table Extraction
PDF OCR & Recognition
OCR PDFs online free — convert scanned PDFs to searchable, extract text via OCR, create searchable PDFs from images. Everything runs in your browser. Also: PDF Editor | Security | Compress | Utilities.
Performance note: OCR downloads ~12MB of language data on first use. Processing is CPU-intensive and happens entirely in your browser. For large documents (>10 pages), expect several minutes of processing. Results depend on scan quality and language.
0%
0%
0%
0%
Scan→Searchable: Renders each PDF page as an image, runs OCR to recognize text, then creates a new PDF with invisible text layer over the page image. The resulting PDF is searchable and selectable.
Image→PDF: Takes one or more images, runs OCR on each, and creates a PDF with each image as a page and recognized text as an invisible overlay. Perfect for photos of documents.
OCR Text: Accepts a PDF or image file. For PDFs, renders each page as an image and runs OCR. Displays recognized text that you can copy. Best for extracting text from scanned documents.
Table Extract: Renders the PDF page as an image, runs OCR on it, and attempts to identify table structures from text positions. Results are approximate — complex tables may not be perfectly reconstructed. For precise tables, use dedicated extraction tools.
What is OCR?
A scanned PDF is essentially a photograph of a page: to the computer it is a grid of pixels, and none of it is text. OCR (Optical Character Recognition) is the technology that looks at those pixels and works out which characters they depict. The result can be plain text you copy out (OCR Text mode) or an invisible text layer placed on top of the page image (Scan→Searchable and Image→PDF modes), which is what makes a scanned document searchable in any PDF viewer.
The one-line version: OCR guesses text from shapes of ink. It is recognition, not reading — the output is a best guess with a confidence score, not the ground truth that a native, digitally-created PDF contains.
The OCR pipeline
Every OCR engine follows roughly the same stages. First the page is rendered to an image (for PDFs, each page is rasterized). Next comes binarization: pixels are classified as ink or paper, so the engine works with black-and-white shapes. The shapes are then segmented into lines, words and individual characters. Finally each candidate is classified — modern engines use trained neural networks that look at whole text lines rather than single letters, which is why they handle blurry or slightly skewed input far better than older single-character matchers. Language matters too: the engine uses language-specific dictionaries and character sets, which is why picking the right language in the selector above improves accuracy.
The invisible text layer
A "searchable PDF" is two things stacked: the original page image you see, plus an invisible text layer positioned behind it. Selecting text in your viewer actually selects that hidden layer. Nothing about the visible scan is altered — the image is untouched, so quality and appearance never change. The trade-off is that copy/paste and search results are only as accurate as the OCR pass underneath. If a word was misrecognized, searching for the correct spelling will not find it, because the hidden layer contains the mistake.
Why DPI decides everything
Recognition accuracy depends on how many pixels each character gets. PDF pages are laid out in points (72 per inch), but a 10 pt character rendered at low resolution collapses into a handful of pixels — at 150 DPI a 10 pt character is only about 21 pixels tall, and at 300 DPI about 42 pixels. Below roughly 20 pixels, serifs close up, letters smear together, and accuracy drops sharply. This is why the long-standing scanning guideline is 300 DPI for OCR: it gives each small character enough pixel detail without wasting file size. Going much higher rarely helps recognition — it mostly makes the scan larger and the OCR slower.
What OCR is not good at
Handwriting, stylized display fonts, and text over busy backgrounds are still weak spots. Tables are a special case: OCR sees characters and their positions, but it does not truly understand rows and columns — Table Extract mode infers structure from alignment, which is why complex tables come out approximate. And if a PDF already contains real text (it was exported from Word, not scanned), you do not need OCR at all: the text can be extracted directly, more accurately, with the PDF export tools.
Common misconceptions
"A searchable PDF is just like a PDF created from Word." No — the text layer is plain recognized text. Copying a paragraph out may lose layout, and any recognition error is baked into the hidden layer.
"OCR recovers the original text." It reconstructs a best guess of it. A 99% accurate page still has roughly one wrong character per 100 — fine for search, risky for legal or numeric content.
"More DPI is always better." Beyond about 400–600 DPI accuracy gains are negligible while file size and processing time keep climbing. 300 DPI is the sweet spot for typical documents.
"Phone photos work as well as scans." Skew, uneven lighting and camera shake all cut accuracy. Photograph the page flat, head-on, and well lit — or use a real scanner for important documents.