← All tools
BROWSER · PDF.js + Tesseract.js OCR (staged local WASM + automatic mixed-script refinement)

PDF to Text (OCR)

Extract text from normal and scanned PDFs. Existing text is preserved and image-only pages use automatic mixed-script OCR.

✓ Reads embedded PDF text✓ OCR fallback for scanned pages✓ Automatic mixed-script detection

Used only on scanned pages. Multiple regions are sampled, so scripts such as Arabic and Chinese can be recognized together.

No PDF selected.

About this tool

Normal PDF text is read directly. Scanned pages are rendered locally, sampled in multiple regions and recognized with the detected writing-system models.

Common uses
pdf to textscanned pdf ocrpdf txtmixed language ocr

How it works

  1. 1Choose or drop a PDF
  2. 2Keep automatic text-source mode or force OCR
  3. 3Keep automatic script detection or open the advanced override
  4. 4Extract and download the UTF-8 text

Frequently asked questions

Related tools