PDF to Word: Text-Based vs Scanned PDFs, and What to Do With a Scan
Whether a PDF converts into editable Word text depends on one thing: whether the PDF contains real text or a picture of text. Here's how to tell in five seconds, and what to do when it's a scan.
Two kinds of PDF that look the same
A text-based PDF is created by software -- Word, Excel, an accounting system, a website's Save as PDF button. Each letter is stored as a real character in a font, so it can be selected, searched and copied.
A scanned PDF is created by a scanner or a phone camera. Each page is a single photograph of paper. It looks identical on screen, but to a computer there are no letters in it, only pixels.
PDF to Word conversion rebuilds a Word document from the text, tables and images in the PDF. With a text-based PDF there is plenty to work with. With a scan, the only thing on the page is a picture.
How to tell which one you have
Open the PDF and try to highlight a sentence with your mouse. If individual words highlight, the PDF has real text. If a whole rectangle is selected, or nothing is, it's a picture.
Or press Ctrl+F (Cmd+F on a Mac) and search for a word you can see on the page. If the reader can't find it, the page is an image.
Our test files showed the difference plainly: the text-based version of our test document contained 801 selectable characters; the scanned version contained none.
What we measured
We made a one-page test document -- a heading, four short clauses and a small rent table -- and converted three versions of it with the same code our tool runs (October 2026). Our PDF to Word tool uses an open-source converter called pdf2docx, with LibreOffice as a fallback.
The text-based PDF converted into an editable Word document: 140 words of text and a real Word table, with "Monthly rent" and "PKR 45,000" in editable cells.
The scanned PDF became a Word document containing one picture of the page and no text.
The third version is the one that catches people out. We ran the scan through Tesseract, a free OCR program, which adds an invisible text layer behind the picture so the PDF becomes searchable. That PDF had 818 selectable characters -- yet our tool still produced one picture and no text. The converter did not use the invisible OCR text at all.
| PDF version | Selectable characters in the PDF | Word result |
|---|---|---|
| Text-based (exported from software) | 801 | Editable: 140 words and 1 Word table |
| Scanned (picture only) | 0 | 1 picture, no editable text |
| Scanned, then made searchable with OCR | 818 (invisible layer) | 1 picture, no editable text |
What to do with a scanned PDF
Ask for the original. If the document came from an office or a company, the Word file or a PDF exported directly from their software will convert far better than any scan.
Use OCR to turn the picture into text. Google Drive offers this for free. According to Google's help page, open drive.google.com on a computer, upload the PDF, right-click it and choose Open with > Google Docs. Google converts the file and opens the result as a Google Doc, which you can download as a Word file with File > Download > Microsoft Word (.docx).
Google lists some requirements for good results: the file should be 2 MB or smaller, text should be at least 10 pixels high, pages must be the right way up, and common fonts such as Arial or Times New Roman work best. Google also notes that bold, italics, font size, font type and line breaks are likely to be kept, but lists, tables, columns, footnotes and endnotes are not likely to be detected -- so expect to rebuild tables by hand.
If a scanned page is sideways or upside down, fix it with our Rotate PDF tool before uploading it to Drive. If the file is over 2 MB, our Compress PDF tool may bring it under the limit, but zoom in afterwards to make sure small text is still sharp, because OCR needs clear letters.
Tesseract is also free and open source if you prefer software on your own computer, but as our test shows, its searchable PDF output still converts to a picture in our tool. Use it for copying text, not as a step before our PDF to Word tool.
Getting the best result from a text-based PDF
Text-based PDFs with a simple layout -- letters, contracts, reports -- convert best. Complex layouts, such as multi-column brochures or text over images, can need tidying in Word afterwards.
Our tool cannot open password-protected PDFs, so remove the password first. Your upload and the Word file are deleted from our server as soon as your download is ready.