How it works
- 1Choose or drop a PDF
- 2Keep automatic text-source mode or force OCR
- 3Keep automatic script detection or open the advanced override
- 4Extract and download the UTF-8 text
Extract text from normal and scanned PDFs. Existing text is preserved and image-only pages use automatic mixed-script OCR.
Used only on scanned pages. Multiple regions are sampled, so scripts such as Arabic and Chinese can be recognized together.
Used only for scanned pages. Select every language present in the document.
No PDF selected.
Normal PDF text is read directly. Scanned pages are rendered locally, sampled in multiple regions and recognized with the detected writing-system models.