PDF to JSON Converter
Keep PDF words, page references and text positions in one JSON file.
Turn selectable PDF text into JSON. Choose reading order, check the extracted words and download. Up to 50 MB and 100 pages.
This is a text export. Pictures, complex tables, equations and the original page design are not rebuilt. Scans need OCR PDF first. For individual pictures, use Extract Images from PDF.
Export a page-based text record from a PDF
Use JSON to review or reuse PDF text in a data workflow. The download groups text blocks by source page, with page dimensions, approximate positions and suggested heading/list types. It includes a title, language code and notices for pages without selectable text.
The output uses adapdf.document.v1. Coordinates are PDF points from the top-left of the visible page. Bounds describe grouped text, not every PDF object or glyph. Extracted values remain strings; account numbers and dates are not guessed as numeric fields.
Choose pages and reading order, then inspect the JSON preview. This is not an invoice field detector, table reconstruction tool or image export. Use PDF to CSV for table data you can review, and OCR PDF before reading scanned pages.
JSON to PDF accepts this format. Document pages lays out the words again; Formatted JSON prints every key and value. Neither option restores the original page geometry. Keep the original PDF with the JSON when source comparison matters.
Best for: People collecting PDF text for data processing or review with source-page references.
How to use PDF to JSON
- Choose a text PDF, select pages and set reading order.
- Create JSON and inspect wording, page blocks and OCR notices against the original.
- Download the JSON. To make a print copy, send it to JSON to PDF and choose the appropriate display mode.
Frequently asked questions
What does the JSON contain?
A schema identifier, title, language, original page count, selected pages, dimensions, text blocks and warnings. Blocks contain text, suggested type and bounds, plus heading level or list type where inferred.
What do the coordinates mean?
They are PDF points from the visible page’s top-left. Rotation and crop affect that frame. Bounds are approximate grouped text positions, not exact glyph shapes or a full page-layout description.
Does it extract invoice totals or named business fields?
No. It exports text blocks, not an invoice-specific schema. Review and map the words yourself. For columns, try PDF to CSV or PDF to Excel.
What if my PDF has two columns?
Choose Two columns for a page divided at its centre, then check that the left column comes before the right. Mixed layouts, text boxes and tables can still need editing. Select pages with the same layout for each run.
Can I convert scanned pages?
Use OCR PDF first, then continue here with its searchable copy. Pages without selectable text are marked in a mixed-document export. A completely scanned selection needs OCR before a download can be created. Check recognized names and numbers.
Are number formats preserved?
Words remain strings, so the exporter does not turn dates, account numbers or decimals into numeric values. Extraction and OCR can still misread characters; verify important values against the source.
Can I convert the JSON back to PDF?
Yes. JSON to PDF prints the JSON source or reflows its page blocks into readable sections. Positions do not reconstruct the original design. JSON exports are limited to 2 MB so they can open in that tool.
Are my documents uploaded?
No. Reading, conversion and previews run in your browser. Document contents are not sent to a conversion service. No account is required.