What this tool does not do
It turns an image into text. It doesn't keep the original layout, and one part of it talks to an outside service.
Plain text only
You get a flat text file, not a searchable PDF with the scan underneath and not a Word document. A desktop tool such as OCRmyPDF adds recognized text back into the original file instead.
Layout does not survive
Columns tend to interleave and tables lose their cells, coming back as a run of numbers. A complex layout needs manual cleanup afterward.
Translate leaves the device
Opening the file, rendering it and running OCR stay local. Clicking Translate text sends the extracted text to MyMemory's API to translate it, after a confirmation prompt. Skip that button for confidential text.
No handwriting
Tesseract is trained on printed text and reads cursive handwriting as confident, wrong text.
Accuracy depends on the scan
Low resolution, uneven lighting, and a rotated or skewed photo all reduce accuracy. A flat, square, well-lit scan reads back best.
English gets a shortcut, other languages don't
The text layer check only runs when the language is English. Pick another OCR language and it always renders the page and runs recognition, even on a PDF that already has selectable text.
Related PDF tools
Frequently asked questions
The file itself is not uploaded. Reading it, rendering its pages and running OCR all happen inside this tab with Tesseract.js and PDF.js. The one exception is the Translate button, which sends the recognized text, not the file, to an outside translation service if you use it.
Only the text currently in the result box, split into pieces under 500 characters, to MyMemory's public translation API over the internet. It asks you to confirm the first time you use it on a page. Your PDF or image is never part of that request.
If you left the language on English and the file is a PDF, the tool first checks for a text layer already in the file, the kind a Word export or a PDF that has already been through OCR carries. When there's enough text there, it returns it directly and skips OCR entirely. Choosing another language always runs full recognition.
No. The result is plain text only, shown in the box and downloadable as .txt. To turn text into a new PDF, use Text to PDF; to keep the original scan and add a hidden text layer, a desktop tool is the better fit.
Recognition reads left to right in blocks, so multi-column pages and tables often come out reordered or flattened into a run of numbers. Low resolution, uneven lighting, a skewed photo, small print or handwriting all lower accuracy too.
Your file stays on your device, mostly
Tesseract.js and PDF.js run in this tab to read the file and recognize its text. Nothing about the file itself is uploaded. If you click Translate text, the recognized text, not the file, is sent to MyMemory's translation API after you confirm.