PDF and image OCR

pdf → txt

Read the text out of a scanned PDF or a photo with OCR that runs in your browser, then copy or download it as plain text.

Settings
QueueNothing is uploaded
Drop a scanned PDF or photo here
Recognizing text...

About this OCR tool

The tool runs Tesseract.js, an OCR engine compiled to WebAssembly, inside a Web Worker. For an English PDF, it first checks whether the file already carries a text layer, the kind a Word export or a PDF that has already been through OCR has, and returns that text right away when there's enough of it. Otherwise it renders each page to an image and recognizes the text on it.

Recognition runs up to three times per page with different image processing: a contrast pass for faint text, an inverted pass for light text on a dark background, and the untouched image as a last resort. Whichever pass reads back the most text is kept.

How to use it

01

Add a file

Drop a scanned PDF or a photo of a page, or choose one with Select file.

02

Pick the language

Choose the language the text is written in from the OCR language list. It defaults to English.

03

Extract the text

Click Extract text. A single page reads back in a few seconds; a multi-page scan takes longer, once per page.

04

Copy or download

Read the result in the text box, then use Copy text or Download .txt. Translate text sends the text elsewhere, so skip it for anything private.

What it does

  • Reads text from a scanned PDF, a photographed page or an image file, one file at a time.
  • Checks an English PDF for an existing text layer first and returns that instantly, without running OCR, when there's enough of it.
  • Retries a page with an inverted or an untouched image if the first pass finds almost nothing, which helps with dark backgrounds and faint scans.
  • Recognizes eleven languages and three English pairs, from Arabic to Chinese, chosen from a dropdown before you convert.
  • Marks each page with a "--- Page N ---" line when a file has more than one page.
  • Lets you copy the result or download it as a .txt file. The Translate button is the only part of this page that sends anything off your device.

Specifications

Specification Details
Accepts One scanned PDF or one image (JPG, PNG and other formats your browser can display) at a time, up to 100 MB
Output Plain text in a text box; Download .txt saves it as ocr_extracted_text.txt
OCR languages English, Spanish, French, German, Hindi, Chinese, Japanese, Portuguese, Italian, Arabic, Russian, and three English pairs
Recognition engine Tesseract.js 5.0.4, running as WebAssembly in a Web Worker, loaded from the cdnjs CDN
PDF handling PDF.js 3.11 reads an existing text layer or draws each page at roughly 144 DPI for OCR
Translate button Sends the extracted text, not the file, to the MyMemory translation API over the internet, after you confirm
Where your file goes Nowhere. It's read and recognized inside this tab.
Good to know

What this tool does not do

It turns an image into text. It doesn't keep the original layout, and one part of it talks to an outside service.

Plain text only

You get a flat text file, not a searchable PDF with the scan underneath and not a Word document. A desktop tool such as OCRmyPDF adds recognized text back into the original file instead.

Layout does not survive

Columns tend to interleave and tables lose their cells, coming back as a run of numbers. A complex layout needs manual cleanup afterward.

Translate leaves the device

Opening the file, rendering it and running OCR stay local. Clicking Translate text sends the extracted text to MyMemory's API to translate it, after a confirmation prompt. Skip that button for confidential text.

No handwriting

Tesseract is trained on printed text and reads cursive handwriting as confident, wrong text.

Accuracy depends on the scan

Low resolution, uneven lighting, and a rotated or skewed photo all reduce accuracy. A flat, square, well-lit scan reads back best.

English gets a shortcut, other languages don't

The text layer check only runs when the language is English. Pick another OCR language and it always renders the page and runs recognition, even on a PDF that already has selectable text.

Related PDF tools

Frequently asked questions

The file itself is not uploaded. Reading it, rendering its pages and running OCR all happen inside this tab with Tesseract.js and PDF.js. The one exception is the Translate button, which sends the recognized text, not the file, to an outside translation service if you use it.

Only the text currently in the result box, split into pieces under 500 characters, to MyMemory's public translation API over the internet. It asks you to confirm the first time you use it on a page. Your PDF or image is never part of that request.

If you left the language on English and the file is a PDF, the tool first checks for a text layer already in the file, the kind a Word export or a PDF that has already been through OCR carries. When there's enough text there, it returns it directly and skips OCR entirely. Choosing another language always runs full recognition.

No. The result is plain text only, shown in the box and downloadable as .txt. To turn text into a new PDF, use Text to PDF; to keep the original scan and add a hidden text layer, a desktop tool is the better fit.

Recognition reads left to right in blocks, so multi-column pages and tables often come out reordered or flattened into a run of numbers. Low resolution, uneven lighting, a skewed photo, small print or handwriting all lower accuracy too.

Your file stays on your device, mostly

Tesseract.js and PDF.js run in this tab to read the file and recognize its text. Nothing about the file itself is uploaded. If you click Translate text, the recognized text, not the file, is sent to MyMemory's translation API after you confirm.