How to Extract Editable Text from Scanned PDFs with Browser OCR
Pull the words out of a scanned PDF or a photograph of a page and get back plain text you can edit, with the recognition running inside your own browser.
The first thing it tries is not OCR
Plenty of PDFs that look like scans are not scans. A file exported from Word, or a scan that some other software has already been through, carries an invisible text layer underneath the picture. Reading that layer gives you the exact characters the document was built from, with no recognition errors at all, and it takes about a second.
So when you hand it a PDF with the language left on English, the tool reads the text layer first. If it finds a reasonable amount of text there, it hands that back immediately and never starts the OCR engine. You will see it finish almost instantly, which is the tell. Pick any other language and it skips this shortcut and goes straight to recognition, so if you have an English PDF, leave the language alone and let it try the fast path.
How a real scan gets read
When there is no text to find, each page renders to an image at 144 pixels per inch and goes to Tesseract, which runs as WebAssembly in your browser. The image gets cleaned up first. The tool measures the lightest and darkest pixels on the page, stretches the contrast between them, and forces every pixel to either black or white, which is what a recognition engine wants to see.
If that first pass comes back with almost nothing, two more run automatically. The second inverts the image, which rescues white text on a dark background, the kind of thing a screenshot of a dark-mode app produces. The third feeds the untouched image straight in, on the theory that the cleanup was the problem. Whichever pass produces the most text wins. You do not choose any of this, and it is why a page that looks hopeless sometimes comes back readable on the third attempt.
What you get back
The result appears in a text box on the page, with a marker line between pages so you can tell where one ended. From there you can copy it or download it as a .txt file. That is the whole output. The tool does not give you back a searchable PDF with an invisible text layer over the original scan, and it does not produce a Word file, so if what you need is a PDF you can search, this gets you the words but not the format.
Layout does not survive either. The text arrives as a flow of lines. A two column page tends to come out with the columns interleaved, and a table loses its cells and becomes a run of numbers. Expect to spend time putting structure back into anything that had a lot of it.
What hurts accuracy
Recognition quality depends almost entirely on the image you feed it:
- Low resolution source scans. A page scanned at 150 DPI has thin, broken letter strokes, and no amount of processing invents the missing pixels back.
- Uneven lighting. The contrast stretch measures the page as a whole rather than region by region, so a photo with a shadow across one side can lose that side entirely to solid black.
- Skew. Even a couple of degrees of rotation from a phone held at an angle costs you accuracy, and nothing here straightens the page for you. Rotate it square before you start.
- Small type, tight leading, and decorative or script fonts. Footnotes and dense legal small print are consistently the worst part of any page.
- Handwriting. Tesseract is trained on printed characters and will produce confident nonsense from cursive.
If you control the scan, the single most useful thing you can do is capture it flat, square and larger. Everything downstream gets easier.
The one button that leaves your device
The recognition is local. Your PDF is read in the tab, the engine and its language data download to you from a public CDN on first use, and no page of your document is sent anywhere. There is one exception on this page and it deserves saying out loud: the Translate button next to the result sends the extracted text to an outside translation service over the internet. The file never goes, but the recognised words do. Skip that button if the document is confidential.
When a desktop tool is the better answer
A tool like OCRmyPDF writes the recognised text back into the original PDF as an invisible layer, so the file still looks like the scan and is now searchable, which is usually what an archive actually needs. Desktop software also deskews and de-speckles before recognition, reconstructs tables, and runs across a folder of hundreds of files unattended. Commercial engines stay ahead on dense, poor quality and multilingual pages. Use the browser when you want the words off one scan, now, and you would rather not upload it to find out what it says.
PDF and image OCR reads the text from your scans in this tab. The file itself is never uploaded.
Open PDF and image OCR