OCR PDF

Make scanned PDFs searchable with OCR.

Recognize text in scanned PDFs using OCR and output a searchable text file. The OCR engine runs fully in your browser — no uploads, no cloud processing.

How to use

01

Select a scanned PDF

Choose the scanned PDF whose text you want to recognize.

02

Choose settings

Pick the render resolution — higher is more accurate but slower.

03

Run OCR

Click OCR. Each page is rendered and recognized with live progress.

04

Download

Download the recognized text as a TXT file.

Features

Fully local OCRThe Tesseract engine runs in your browser via WebAssembly — no cloud API.
Live progressSee per-page recognition progress.
Resolution controlBalance accuracy against processing time.
English modelHigh-quality English language data bundled with the site.

Privacy & security

OCR runs entirely in your browser. Scanned pages are never transmitted to any server.

Limitations

  • OCR accuracy depends on scan quality, fonts and layout — handwriting and low-quality scans work poorly.
  • The first run loads the language model (~10 MB) from the site's own assets; processing is CPU-intensive.
  • Only the English language model is currently bundled.

Frequently asked questions

Yes — the Tesseract engine and language data run locally via WebAssembly. Nothing is uploaded.

Recognition is CPU-heavy and runs on your device. Higher DPI settings increase accuracy and time.

Currently English only; additional language models may be added later.

Yes — free, no account, no limits.