How to extract text from a scanned PDF with OCR, free in your browser
Published on September 18, 2026
To extract text from a scanned PDF, use OCR: pick the language, click "Recognize and download" and get a .txt with the text. Free, in your browser, with no upload.
How to extract text from a scanned PDF
A scanned PDF is a photo of the page, so the text has to be recognized, not copied. AtlasDocs OCR reads each page in your browser, identifies the letters and returns a .txt file with the text of every page, ready to copy, search or paste into Word. The original PDF is not changed: you get the text in a separate file.
- Open the OCR tool and upload the scanned PDF or the photo of the document (JPG, PNG, WebP or GIF). You can upload several files at once.
- Pick the document's language in the menu by the button (it defaults to Português): Português, English or Español. The wrong language mixes up accents and confuses similar-looking letters.
- Click "Recognize and download" (with several files, the button reads "Recognize 3 files"). The message "Recognizing… (may take a while)" shows while your browser reads page by page.
- You get one .txt file per document, named after the original. The text follows the page order, with no marker where each page starts; open it in Notepad or Word and review it.
The first time, your browser downloads the reading program and the model for the chosen language (a few megabytes) before starting; after that they're stored on your device and recognition starts much faster. Switching languages downloads that language's model the first time. Long PDFs take a few minutes, because each page is turned into an enlarged image and read one at a time. If nothing is recognized, the tool shows a red warning that no text was found: check that the scan is sharp and the language is right.
What OCR is and why you need it
OCR stands for optical character recognition. It's what lets a computer look at an image of text and work out which letters are there. To a computer, a scanned PDF or a photo of a document is just a picture: there are no words inside it. OCR walks through the image, identifies each character and builds real text that you can select, copy and edit in any program.
How to tell if a PDF is scanned (and which tool to use)
There are two kinds of PDF that look identical on screen. The first was born digital (exported from Word, say) and already holds real text inside: you can select a word with your cursor. The second is an image, from a scanner or a photo, and nothing is selectable, no matter how sharp it looks. This second case is where OCR comes in: it reads the image and hands back the text in a separate .txt file. If what you want is the document in Word, with the page's formatting and not just the text, see how to convert PDF to Word keeping the formatting.
- If you can click and drag to select a single word, the PDF already has text. Don't use OCR: the Extract text from PDF tool saves all the text to a .txt at once, with no letter-by-letter guessing, in seconds. If it warns that the PDF has no extractable text, it's a scanned document: then OCR is the way.
- If you try to select text and the whole page gets highlighted as one block, it's a scanned document and OCR is the way.
- Photos of documents taken with your phone (JPG, PNG, WebP, GIF) are images too and work with OCR.
- HEIC (the iPhone photo format), TIFF, BMP and AVIF are not accepted by the OCR tool: convert them to JPG or PNG first, then upload.
If the photo came from an iPhone as HEIC, convert it first with the HEIC to JPG tool, which also runs in your browser; TIFF or AVIF images go through the TIFF/AVIF to JPG tool, and BMP through the Convert BMP to JPG, PNG or WebP tool. Note: from a multi-page TIFF, like the ones some scanners produce, only the first page is converted. Then upload the resulting JPG to OCR.
HEIC to JPG →TIFF/AVIF to JPG →Does OCR keep the formatting or make a searchable PDF?
- It does not produce a searchable PDF: the original PDF stays as it was, with no text you can search. If you need a PDF with the text, turn the .txt into one with the Convert text to PDF tool; the new file holds only the text, not the look of the scanned page.
- It does not keep formatting: bold, columns and tables become plain running text. Numbers and proper names are where most mistakes slip through; check them afterward.
- Input: PDF and .jpg, .jpeg, .png, .webp and .gif images. Several files at a time.
- Output: one plain-text file (.txt) for each file you upload, with the text of every page in document order.
- Languages: Português, English and Español, one per run. Document in two languages: run it twice, one language each time, and keep the best passage from each.
Why OCR gets letters wrong and how to get better results
- Use sharp, well-lit scans whenever you can; shadow, blur and glare get in the way of reading.
- Keep the document as straight as possible: crooked pages confuse line detection.
- Pick the right language before clicking: the selector decides how accents and special characters are interpreted.
- Double-check names, dates and numbers after downloading; those are the details where mistakes most often slip through, especially in shaky photos.
AtlasDocs OCR runs inside your browser: your document is never uploaded to any server. The only downloads are the tool's own parts (the reading program and the language model), which stay on your device. Ideal for papers with personal data.
