Gearyforgemore tools

Merged cells in PDF tables

Spreadsheets flattened from PDFs usually break on merged cells: the grouped header or the label covering five rows becomes one value followed by a row of gaps. gridlift keeps the merge. A cell that spans rows or columns is stored with its value and the size of its span, the preview shows it spanning exactly as the PDF prints it, and at export you decide what the covered positions should contain — repeat the value into each one, or leave them empty.

Where merges come from

PDF authors merge cells for two common reasons. Grouped headers stack a label like a year above the months beneath it, so the year cell spans several columns. Grouped rows — regions above their cities, categories above their items — use one label cell spanning several rows. Both shapes are ordinary in invoices, schedules, and financial statements, and both become garbage in a naive grid flatten.

How a merge survives detection

When detection finds a cell that spans several rows or columns, it keeps the value in its original position and records how far the span reaches in each direction. The table is therefore stored as a grid with spanning cells, not as loose text, which is what makes an honest export possible at all.

The preview shows the merge as printed

In the preview, a spanning cell renders across its full range, and the positions it covers are not rendered as separate empty cells. What you see is the table as the PDF draws it. This does not change when you switch the export policy — the choice applies to the exported file, not to the preview.

Repeat or empty: you choose at export

Two policies cover the exported file. Repeat merged-cell values fills every covered position with the span's value, so each row stands on its own — what pivot tables, filters, and databases want. Leave merged cells empty keeps the covered positions blank, closer to the printed layout. Repeat is the default, and the policy applies to JSON exports as well as CSV.

Frequently asked questions

Does repeating a value change my data?

No. The value itself is untouched; repeating only copies it into the positions the span covers. Everything else in the table is exported exactly as detected.

Which policy should I pick?

Repeat when the CSV feeds analysis — filters, pivots, and databases need every row populated. Empty when you want the file to mirror the printed layout, or when a downstream tool reads sparse tables.

Does switching the policy change the preview?

No. The preview always shows the merge as printed in the PDF. The policy takes effect in the downloaded file.

Does the policy apply to JSON too?

Yes. JSON exports take the same merged-cell policy into account, so a repeated export and an empty export differ the same way in either format.

What if the PDF only looks merged?

If the PDF prints the value once, detection records one spanning cell. If it prints the value in every row, those are ordinary cells with repeated text, and they export as they are. The policy only governs real spans.

Related pages

  • PDF to CSV — Extract tables from a PDF to CSV in your browser. Every table on every page is detected automatically and previewed free; the CSV download comes with a one-time unlock.
  • PDF to JSON — Turn the tables in a PDF into JSON - one selected table or every table in a single document - with each table addressable by id, page and source. Runs entirely in your browser.
  • Multi-page PDFs — gridlift scans every page of a multi-page PDF and lists each table with its page, size and detector - keeping page boundaries clear instead of stitching a guess.

Open gridlift to detect the tables in your PDF and preview them free.