PDF to Excel: Why Some Tables Don't Convert Cleanly (With Test Results)
A PDF doesn't actually contain a table -- just text and lines placed on a page. How well a table converts to Excel depends on how clearly those lines mark out the cells. Here's what we found.
Why tables are hard to get out of a PDF
A spreadsheet knows which cell every value is in. A PDF doesn't: it records each piece of text and each line as a drawing instruction at a position on the page. When you look at a PDF table you see rows and columns, but the file only knows that some text sits near some lines.
So a PDF to Excel converter has to rebuild the table. Our tool uses an open-source library called pdfplumber to find table structures on each page, mainly from the ruling lines drawn around and between cells, and then writes each table into a worksheet, with the text found above and below it.
How our tool lays out the Excel file
Each page of the PDF that has content becomes its own worksheet, named Page 1, Page 2 and so on.
Text around a table is kept. A heading above a table and a total below it are written into the same sheet, one line per row, so nothing on the page is silently dropped.
A page where no table is detected still contributes its text, one line per row in column A.
Every value is written as text. The tool does not guess whether 1,000.50 is a number, a code or a date, so it leaves that decision to you. The next sections show how to turn those values into real numbers.
What we measured
We built five test PDFs of a typical invoice table -- Date, Description, Qty, Unit price and Amount -- and converted each with the same code our tool runs (October 2026).
| Test PDF | Result in Excel | What to do |
|---|---|---|
| Table with ruled lines around every cell | All 5 columns correct; heading and Total line kept | Convert number columns from text to numbers |
| Same table with no lines (borderless) | Each whole row placed as one line in column A | Use Data > Text to Columns, then check every row |
| Table with a merged title cell across the top | Title placed in the first column only; table below correct | Merge or re-centre the title in Excel if needed |
| Ruled table running across 2 pages | Two sheets (Page 1 and Page 2); header row repeated on sheet 2 | Copy the rows into one sheet and delete the repeated header |
| Scanned PDF (a picture of the same table) | Not converted - message that no text or tables could be found | Needs OCR first - see below |
Fix 1: numbers stored as text
In our test, values like 1,000.50 and PKR 1,000 arrived in Excel as text. You'll often see a small green triangle in the cell's corner, and SUM formulas will ignore those cells.
For plain numbers, select the column, click the warning icon that appears next to the selection and choose Convert to Number.
For values with a currency label or other text, remove the text first. Select the column, press Ctrl+H, type the label exactly as it appears -- for example PKR followed by a space -- in Find what, leave Replace with empty, and choose Replace All. Then use Convert to Number.
Fix 2: borderless tables land in column A
Without ruling lines there is nothing reliable to show where one column ends and the next begins, so in our test each row came through as a single line of text in column A.
Excel can split those lines for you. Select column A, choose Data > Text to Columns, pick Delimited and tick Space. Check the result carefully: a value that itself contains spaces, such as a description like "Office chair", will be split across two columns and needs to be joined back together. If you can get the original file the PDF was made from, such as the Excel or Word file, use that instead.
Fix 3: one table spread over several sheets
Because each PDF page becomes its own sheet, a long table is split into pieces -- and if the PDF repeats the header row on every page, each piece starts with that header. Copy the rows from Page 2 onwards onto the end of the Page 1 sheet, then delete the extra header rows. Sorting or filtering by the first column can help you spot the repeats.
Scanned PDFs need OCR first
A scanned PDF is a picture of a page. There is no text in it for a converter to read, so our tool stops with a message that no text or tables could be found. Our site does not currently offer OCR (optical character recognition, which turns a picture of text into real text).
To check whether a PDF is scanned, open it and try to select a word, or search for a word with Ctrl+F. If you can't, it's a picture. The best option is usually to ask the sender for the original file, or for a PDF exported directly from the software that created it.
Getting the best result
PDFs exported from accounting or banking software, with lines around every cell, convert best. Check totals in Excel against the PDF after converting, convert number columns from text, and keep the original PDF for reference. Your upload and the Excel file are deleted from our server as soon as your download is ready.