Parse PDF: Turn a Document into Structured Data

Publish date
Sep 4, 2026
AI summary
Parse PDF into structured data by extracting headings, paragraphs, tables, and figures; use OCR for scans, choose standard or advanced quality, and follow API or browser paths, verifying output and selecting parse, extract, or OCR based on needs.
Language
People type parse pdf when they need the document as data: headings, paragraphs, tables, and figures they can feed into code, a sheet, or another model. They do not mean “open it in a viewer.” They also do not mean installing a random npm package with a similar name.
This page is that job. It sits under the Extract pillar: Extract Data from PDF: Get Fields and Tables Out. Parse is a cluster under Extract, not its own pillar. If the file is a scan with no text layer yet, start with OCR PDF: Make a Scanned File Searchable.

What parse returns

Parse turns a PDF into structured content: markdown plus a list of content blocks (and related image data) you can use downstream.
On PDF.ai, the live parser / parse-pdf pages describe the same idea: parse PDFs with OCR and layout detection, and get structured JSON with headings, paragraphs, tables, and figures.
After a clean parse, three things should be true:
  • You have structure, not only a flat text dump
  • Tables are usable as data, not only as a screenshot
  • You can point at a block and know which page it came from
Practical rule: parse is finished when the layout is data you can use. A searchable PDF alone is OCR. A handful of named fields is extract.

Parse vs OCR vs extract vs chat

Important distinction: these jobs start from the same file and answer different questions.
Job
What you get
What you still do not have
OCR
Readable / searchable text (text layer or OCR output)
A full layout schema, or named business fields
Parse
Structured blocks: headings, paragraphs, tables, figures
A small custom schema of only the fields you named
Extract
Named fields you defined (plus citations when returned)
A rewrite of the whole document
Chat
Answers in sentences
A stable schema for every file in a folder
Extract for named fields lives at https://pdf.ai/extract-pdf. The Extract pillar explains when to stop at fields instead of full parse: Extract Data from PDF: Get Fields and Tables Out.
Practical rule: if you need invoice_number + date + total, use extract. If you need the document tree for RAG or table rows as layout, use parse.

Not the npm package

Search results for “parse pdf” often mix in library names (for example npm packages). This guide is about the document job: turn a PDF into structured data you can use. It is not a package README and it will not walk through installing someone else’s module.
If your real need is “call an HTTP API from my backend,” keep reading. If your real need is “ship a searchable file once,” use the OCR browser path instead.

How to parse a PDF in practice

You can try a file in the browser, or call the API.

1. Confirm the file can be read

Open the PDF and try to select a word. If nothing highlights, treat it as a scan. OCR / standard parse quality is built for that case, but a blurry phone photo still fails. Fix the capture first when you can.

2. Browser path (no key for a quick try)

Go to https://pdf.ai/parser or https://pdf.ai/parse-pdf. The product copy says you can try a file and get structured output. Upload, wait, then inspect tables and text blocks before you trust them.
notion image

3. API path

Endpoint: POST https://pdf.ai/api/v2/parse with header X-API-Key.
From the docs:
  • Send file, url, or docId
  • quality: standard (OCR; supports lang_list) or advanced (VLM)
  • llm: optional image LLM processing
  • Cached results with the same settings use 0 credits
Credit examples published on the docs:
  • Parse Standard (OCR): 1 credit/page
  • Parse Advanced (VLM): 2 credits/page
  • Free API plan: 200 credits/month at $0, no credit card required
Example shape (Python), matching the docs:
import requests

url = "https://pdf.ai/api/v2/parse"
headers = {"X-API-Key": "YOUR_API_KEY"}

with open("/path/to/document.pdf", "rb") as f:
    files = {"file": f}
    data = {
        "quality": "standard",
        "lang_list": '["en"]'
    }
    response = requests.post(url, headers=headers, files=files, data=data)

print(response.json())
A successful body includes markdown, contents, images, pageCount, and docId (see the Parse docs). Keep docId when you will call extract or split later; the docs say caching helps those follow-on calls.

4. Check the output

  • Spot-check the first heading and a mid-document paragraph
  • For a table: first row, a middle row, and the last row you care about
  • Confirm page numbers on blocks you will store
Practical rule: structured JSON can still be wrong on one cell. Read the source page before you write to a database.

When parse goes wrong

  • You wanted five named fields. Use extract, not a full parse dump.
  • The table is a photograph. Improve the scan, then parse again.
  • You expected a Node library tutorial. Wrong article. This page is the document-parse job and the PDF.ai parse API / playground.

A short action plan

  1. Name the output: full layout vs a short field list.
  1. If you only need fields, go to extract. If you need structure, stay on parse.
  1. Try https://pdf.ai/parser on one file, or call POST https://pdf.ai/api/v2/parse.
  1. Verify headings and table rows against the PDF.
  1. Keep the original file until the values you will use are correct.
For named fields and tables as a product job, keep the Extract pillar open: Extract Data from PDF: Get Fields and Tables Out. For scans first, use the OCR pillar: OCR PDF: Make a Scanned File Searchable.
If you want the rest of the PDF workflow after the structure is out, start at PDF.ai.