🔍
PDF/PDF OCR

PDF OCR

Turn scanned PDFs into selectable text with Tesseract — no uploads, no accounts, no watermarks.

100% LOCALTesseract WASMPage rangeEditable output100% local
🔍
Drop a scanned PDF here, or click to browse
English OCR · runs in your browser
toof

Drop a scanned PDF and turn its pages back into text

Data Source & Legal Disclaimer
Effective: 2026Last updated: TodayUpdate: Manual review
Sources: Tesseract.js (Tesseract 5 WASM)

OCR runs entirely in your browser via Tesseract WASM. On first use the English language model (~15 MB) is downloaded once from a public CDN and cached; your PDF itself is never uploaded.

See all data sources & update policy →

Why this OCR is different — and what it honestly can't do

Every page is rasterized at 2× zoom first, because Tesseract accuracy drops sharply on 72 dpi input — the same scan that looks fine on screen is often too coarse to read reliably. The English model (~15 MB) is fetched once on your first run and cached by the browser; from the second document onward OCR starts instantly. Only English is supported — other languages are a real roadmap item, not a hidden limitation. Pages with no recognizable text come back as an explicit [no text recognized] marker instead of an empty gap, so you always know what the engine did.

Scan → 2× raster → Tesseract WASM → editable text
Scanned PDFimages, no text layerCanvas 2× rastersharpens for OCRTesseract WASMin your browser tabEditable textcopy or .txt~15 MB English model, downloaded once, cached forever · page markers between pages

All in the browser: the PDF never uploads, and the recognition model is a one-time ~15 MB download.

Digitizing a 12-page faxed contract — Priya's archive

Priya received a signed contract as a flat scan. Copy-paste returns nothing because there is no text layer underneath the image.

  1. Load the scan:The tool reads the page count (12) and pre-fills the range 1–12; she narrows it to pages 3–9 where the amendment lives
  2. First-run model:The first run downloads the 15 MB English model once — the progress bar says so before it starts, and the next document skips this step
  3. Page-by-page OCR:Each page renders at 2× and is recognized in turn; pages already carrying a real text layer pass through in seconds, photographic scans take longer
  4. Fix and keep:She corrects two misread names directly in the editable output, copies it into the archive system, and downloads contract_ocr.txt as backup
↩ Back to OCR

About this tool

What is this tool?

OCR scanned PDFs online free with Tesseract. Page ranges, live progress, editable text output.

Tesseract WASMPage rangeEditable output100% local

Why 2x Rasterization Before Recognition

Tesseract accuracy drops sharply on coarse input, and the same scan that looks fine on screen is often only 72 dpi - too few pixels for reliable character shapes. Rasterizing every page at 2x zoom before recognition materially sharpens the input, which is why this tool renders first and recognizes second rather than feeding Tesseract the raw page.

Local Model, Honest Scope

The English model (~15 MB) downloads once on your first run, caches in the browser, and every later document starts instantly - all recognition happens locally, and your PDF never uploads. Pages with no recognizable text come back as explicit no-text-recognized markers instead of silent gaps, the page range lets you skip non-scan sections of hybrid documents, and the editable output means the two misread names in a faxed contract are a five-second fix.

Frequently Asked Questions

Why is the first run slower?

On first use the tool downloads the ~15 MB Tesseract English model from a public CDN and caches it in your browser. Every later run starts instantly. The PDF itself is never uploaded - only the one-time model download ever leaves the page.

Which languages are supported?

English only, honestly. Other languages are a roadmap item, not a hidden limitation - the tool would rather say so than silently return garbage on a German or Spanish scan.

How accurate is the OCR?

Accuracy depends on scan quality: clean 300 dpi documents recognize reliably, while faxed or photographed pages will show errors on small text and unusual fonts. The 2x rasterization before recognition materially improves results over raw page images, and the output is editable so corrections take seconds.

Related tools

Joke of the Day
Sep 7

Why did the elephant paint its toenails red?

100% Free, Forever

Keep Tools Free for Everyone

No paywalls, no signups, no data sold. Built by a solo developer who believes useful tools should be accessible to everyone.

Support me on Ko-fi— keep tools free

100% of proceeds go towards hosting & building more free tools.