Docs · guide

PDF to Excel: why it is the hardest one

By Alberto Gulotta · Updated · 9 min read

pdf to excel is the conversion people complain about most, and the complaints are justified. It is genuinely harder than every other job in this section, for a reason that has nothing to do with the quality of the converter.

Why a table in a PDF is not a table A spreadsheet stores cells with rows and columns. A PDF stores marks at coordinates: the grid you can see is drawn lines, and the numbers are text placed near them. Converting back means guessing where the cells were. In a spreadsheet In a PDF Every value lives in a cell that knows its row and column. The same marks, at fixed positions. No rows, no columns, no cells.
This is why converting a PDF back into a spreadsheet is the hardest job in this section: the structure was thrown away when the file was made, and every converter is guessing it back.
Excel will attempt it itself, and Microsoft says where the attempt breaks

The route is in the ribbon rather than on a website: Data > Get Data > From File > From PDF. Microsoft’s description of what happens next is the honest part — “in Navigator, select the file information you want, then either select Load to load the data or Transform Data to continue transforming the data in Power Query Editor.” Navigator shows what the connector believes it found. It is a list of candidates, not a list of tables that were ever really there.

Then Microsoft documents the failure this page is about, in its own words: “in cases where multi-line rows aren’t properly identified, you might be able to clean up the data using UI operations or custom M code” — with Table.FillDown offered as the repair for “misaligned data”. A row whose text wrapped onto a second line is exactly the case where marks-at- coordinates stops looking like a grid, and the connector’s own manual expects you to be fixing it afterwards rather than getting it right first time.

Two more limits are published for big documents. “Try selecting pages one at a time or one small range at a time using the StartPage or EndPage options, iterating over the entire document as needed.” And for the case that sounds easiest but is not: “if the PDF document is one single, huge table, the MultiPageTables option can be collecting very large intermediate values, so disabling it might help.” A hundred-page report is not converted in one gesture; it is converted in slices, by someone watching the result.

The clearest description of the problem comes from a tool built for it. Tabula is free and open source, made with the support of ProPublica, La Nación DATA, Knight-Mozilla OpenNews and The New York Times, and it offers two extraction modes — which is the whole argument of this page in one interface. The documentation of tabula-java, the library that does the extracting, names them by what each one assumes: lattice is the mode to force “if there are ruling lines separating each cell, as in a PDF of an Excel spreadsheet”, stream the one to force “if there are no ruling lines separating each cell”. A program that has to be told which kind of picture it is looking at is not reading a table, it is inferring one. Tabula also states the limit that decides whether any of this is worth attempting: it “only works on text-based PDFs, not scanned documents”. If the text cannot be selected with a cursor, the file is a photograph of a table and belongs on the OCR guide instead.

One line in the connector documentation saves a step for anyone trying to extract a table from a PDF that lives on a website: “if the PDF file is online, use the Web connector to connect to the file.” The document does not have to be downloaded first, and the query keeps pointing at the address rather than at a copy in the Downloads folder — so refreshing it later fetches whatever is published now, instead of whatever was published on the day you saved it.

A spreadsheet is a grid of cells that know their own row and column. A PDF is a set of marks at fixed positions on a page. When a table is exported to PDF, the grid is drawn as lines and the values are placed as text near them — and the structure that made it a table is discarded. Converting back means looking at the positions and guessing where the cells used to be.

The searches show the two halves clearly. Pdf to excel and pdf to csv are people trying to get data out. Excel to pdf and convert excel to pdf are people trying to send a sheet to somebody who should not edit it — a far easier job, free in every program. Excel online is somebody without the software installed, and csv file, csv to excel, open csv and xlsx are people wrestling with the formats themselves, usually after something arrived looking wrong.

That is why the same file can convert perfectly in one tool and produce nonsense in another, and why a bordered table converts better than a tidy one with no lines at all.

Where to start

Five ways in. The last one is the answer nobody wants and everybody should hear.

“I need the numbers out of this PDF.”
Start at getting data out
“The columns came out merged.”
That is why it goes wrong
“I need to send a spreadsheet as a PDF.”
That is the easy direction — the other direction
“It is a CSV and Excel is mangling it.”
Go to csv files
“Is there a way that always works?”
Yes, and it is in getting data out

Getting data out

From a PDF into something you can calculate with. Ordered by how reliable each route is, which is not the order the search results suggest.

Why it goes wrong

The failures are predictable once you know what the converter is looking at. Each guide here names a cause and what to change.

The other direction

Excel to PDF is the easy half, free everywhere, and mostly about page setup rather than conversion.

CSV files

The plain-text format underneath all of this. Small, universal, and responsible for more quiet data damage than any other file type here.

Not covered here. It will not tell you a converter will get your table right. Some tables convert perfectly and some cannot be converted at all, and which one you have is decided by how the PDF was made — not by which product you choose.

It will not rank the pdf to excel tools yet. That verdict needs the same set of tables — bordered, borderless, multi-page and scanned — run through each of them with the results published.

And it will keep saying the unglamorous thing: ask for the original file. It is one message, it always works, and it is faster than every route below it. A page that skips that advice in order to recommend a tool is not helping you. What holds instead is simple: the most reliable route on this guide costs nothing and is a single message.

Step 1 of 21 in half an hour in a spreadsheet

The order the work actually happens in: the data arrives, it gets cleaned, it gets counted, it gets a chart, and then somebody else opens it.

After this
How to remove duplicates in Excel, and what counts as one

Turning a file into the format they asked for

The same job in six directions. Which one you need depends on what arrived.

The same job, in the other places it comes up
How to turn a picture into a PDF, without uploading it
How to convert a PDF to JPG, and when not to
How to convert PDF to Word, free, without uploading it
Word to PDF: the route that is already in Word
How to convert PowerPoint to PDF, and why you should anyway

Sources

  1. Microsoft, PDF connector — Power Query, Microsoft Learn (the route to connect to a PDF file and choose what to load in Navigator, with Load or Transform Data; the note that in cases where multi-line rows are not properly identified the data may need cleaning up with UI operations or custom M code, using Table.FillDown for misaligned data; and the strategies for large files, namely selecting pages one at a time or a small range using the StartPage or EndPage options, and disabling MultiPageTables when the document is one single huge table) — learn.microsoft.com, read 24 August 2026.
  2. Microsoft on saving or converting an Office file to PDF, what PDF preserves, and its warning that a PDF does not prevent editing — support.microsoft.com, read 6 September 2026.
  3. Tabula — the free tool for extracting tables from PDFs, made by Manuel Aristarán, Mike Tigas and Jeremy B. Merrill with the support of ProPublica, La Nación DATA, Knight-Mozilla OpenNews and The New York Times, for its limit: it works only on text-based PDFs, not on scans — tabula.technology, read 16 September 2026.
  4. tabula-java — the library and command-line interface that powers Tabula, for the two extraction modes and what each one assumes: lattice “if there are ruling lines separating each cell, as in a PDF of an Excel spreadsheet”, stream “if there are no ruling lines separating each cell” — github.com, read 16 September 2026.

Written by Alberto Gulotta

Founder and editor of AI Tools Primer, writing from Palermo, Italy. Thirty-five years of taking computers apart, starting with a Commodore 64 — the long version is on the about page.

Something wrong on this page? Write to aitoolsprimer@gmail.com and it gets fixed.

Written on 21 August 2026 · last checked 16 September 2026.

Independence and limits

No affiliate links and no paid placements anywhere on this site. Nobody pays to appear here, and no company has seen this page before you did.

This is general information, not professional advice. Where a page touches money, health, safety or the law, it names its source and the date it was read — and your situation may still differ. See the privacy page and the cookie policy.