PDF tables to JSON
gridlift exports the tables it finds in a PDF as JSON, in the shape integration code wants: every table is an object carrying an id, the page it sits on, the detector that read it, its column names, and its rows as objects keyed by those columns. Export just the table you selected, or every table in the document in one file. Like everything in gridlift, it runs entirely in your browser - the PDF is never uploaded - and JSON downloads are part of the same one-time unlock as CSV.
Two scopes: one table or all of them
The export panel offers a JSON scope. Selected table writes a single table object for the table you picked in the list. All tables writes a document holding your PDF's file name, a tableCount, and a tables array with every detected table in reading order. For a three-table statement, that is one file with three addressable tables rather than three exports to stitch together afterwards.
The shape of a table object
Each table carries: id, like "p1-t1" (page 1, first table on that page); page, the PDF page number; source, either "lattice" for a table read from ruled lines or "stream" for one read from whitespace-aligned text; columns, the header names; and rows, an array of objects whose keys are those column names. Nothing in the file is ambiguous about where a row came from.
Headers become object keys
Column names come from the header rows gridlift detects, at most three. Nested headers are flattened into readable paths joined with " > ", so a stacked header becomes "Year > January" instead of two cells with the same name. Duplicate column names receive a numbered suffix - "(2)", then "(3)" - and a column with no header text falls back to "column_N", so every key in every row object is unique.
Values arrive as strings
Every value in rows is a string, exactly as the text appears in the PDF once whitespace is normalized - a currency figure arrives as "1,234.50", not as a number with silent assumptions about thousands separators or locales. Convert deliberately in your own code, where a failed parse is loud instead of quiet.
Frequently asked questions
Is my PDF uploaded to a server?
No. The PDF is read and processed in your browser, and the file itself never leaves it. If you choose to pay, a Stripe-hosted checkout page opens in a new browser tab.
How does the merged-cell policy affect the JSON?
The same way as CSV. Repeat fills every position a span covers before rows are written, so each row object is complete; Leave empty keeps covered positions as empty strings. Switch the policy in the export panel and download again.
What do id and page mean?
id is p{n}-t{k}: page n, table k in reading order on that page. page is the PDF page number. Together they make every table addressable, so downstream code can rely on the document's structure instead of counting.
What does the source field tell me?
How the table was read: "lattice" means it was reconstructed from ruled lines, "stream" means it was recovered from whitespace-aligned text with no ruling. It is a hint about how much to trust a table's geometry.
Can I get the JSON without the unlock?
No. Like full rows and CSV downloads, JSON export is part of the one-time unlock. Detection and the ten-row preview of any table are free, so you can verify the shape of the export before you decide to pay.
Related pages
- PDF to CSV — Extract tables from a PDF to CSV in your browser. Every table on every page is detected automatically and previewed free; the CSV download comes with a one-time unlock.
- Merged cells — PDF tables with merged cells export without broken rows: the preview shows each span as printed, and you choose whether covered cells repeat the value or stay empty.
- Multi-page PDFs — gridlift scans every page of a multi-page PDF and lists each table with its page, size and detector - keeping page boundaries clear instead of stitching a guess.
Open gridlift to detect the tables in your PDF and preview them free.