Gearyforgemore tools

Tables in multi-page PDFs

Long documents are where table extraction usually falls apart: detectors stitch tables across page breaks, guess at headers carried over, and hand you one merged blob. gridlift takes the opposite approach. Detection is page-aware: every page is scanned in order, each page's tables are detected on their own, and the list shows which page each table came from. Nothing is merged across a page break on a guess - which is exactly what you want when page 4 repeats the layout of page 3 with different numbers.

Every page is scanned, in order

gridlift reads the PDF page by page and runs detection on each page independently. The table list follows the document: page 1's tables first, then page 2's, each page's tables ordered top downward. A 40-page report and a single invoice take the same path - no configuration, no page selection.

Every table keeps its page identity

Tables are labelled p{n}-t{k}: page n, table k on that page. In the list, each entry shows the page number, the dimensions (rows by columns) and which detector read it - ruled-line or whitespace. When you export all tables to JSON, every table in the document carries its id and page, so downstream code can address, say, the second table on page 7 without counting.

Repeated layouts stay separate tables

Invoices, pay slips and statements often repeat the same table layout across pages with different data. gridlift treats each page's table as its own table rather than stitching them into one. That is deliberate: stitching guesses where one page's table ends and the next begins, and a wrong guess corrupts rows in ways that are easy to miss. If you want the pages combined, do it downstream where you can see the join.

Getting everything out of a long document

A CSV download exports the table you have selected, one file per table. To take every table in the document at once, export all tables to JSON - a single file in which every table is an object with its id, page and rows. Detection and the ten-row preview of any table are free on documents of any length; full rows and downloads come with the one-time unlock.

Frequently asked questions

Is there a page limit?

No. gridlift detects tables on every page of the document, however long it is, and the ten-row preview of any table is free. Seeing every row and downloading CSV or JSON is part of the one-time unlock.

Will a table continued across a page break be stitched together?

No. Pages are detected separately, so page boundaries remain clear, and each page's tables are listed on their own. If a table genuinely continues, export both pages' tables and combine them in your spreadsheet, where you can check the join.

How do I know which page a table came from?

The table list shows the page number with each entry, and the id encodes it - p3-t2 is the second table on page 3. The JSON export keeps the page number on every table as well.

Can I download all tables as one CSV?

Not as one file - CSV export is per selected table. To get every table in a single file, use the all-tables JSON export; from there it is a short script or a spreadsheet import to flatten however you like.

Does all of this happen in my browser?

Yes. The PDF is read and processed in your browser and never uploaded, no matter how many pages it has.

Related pages

  • PDF to CSV — Extract tables from a PDF to CSV in your browser. Every table on every page is detected automatically and previewed free; the CSV download comes with a one-time unlock.
  • PDF to JSON — Turn the tables in a PDF into JSON - one selected table or every table in a single document - with each table addressable by id, page and source. Runs entirely in your browser.
  • Merged cells — PDF tables with merged cells export without broken rows: the preview shows each span as printed, and you choose whether covered cells repeat the value or stay empty.

Open gridlift to detect the tables in your PDF and preview them free.