Extract Data from PDF: Get Fields and Tables Out
Publish date
Sep 7, 2026
AI summary
The guide explains how to extract structured data—named fields and tables—from PDFs, distinguishing this task from OCR (which adds a searchable text layer) and chat (which provides answers). It outlines practical steps: verify the PDF is selectable, use OCR if needed, define the required fields or tables, upload the file to the appropriate extraction tool, and validate the output against the source. It also offers tips on when to choose extract, OCR, or chat, and provides an action plan for efficiently obtaining reliable data for spreadsheets or databases.
Language
You can see the invoice number on the page. The date, the total, and a table of line items are right there. You still cannot drop those values into a sheet without retyping them.
Search can find a word. Chat can answer a question. Neither gives you a row of fields you can paste into a spreadsheet.
Extract data from PDF is the job of pulling structured data out of the file: named fields and tables. This guide stays on that job. It is not about making a scanned file searchable, chatting with a document, or writing extraction code. It is about getting the values you need, then checking them against the page.
What extract actually returns
When people search extract data from PDF, they usually want one of two outputs:
- Fields. An invoice number, a due date, a total, a vendor name, a contract party, an amount.
- Tables. Line items, a rate card, a statement of hours. Rows and columns, not a screenshot of a grid.
PDF.ai describes the product in those terms: turn PDFs into structured data, parse a document, and extract custom fields. Extract is the name of this pillar because that is the phrase people search. Parse is a related job under the same roof. It is not a different pillar.
After a clean extract, three things should be true:
- You have named values, not a blob of copied text.
- Tables come out as rows, not a paragraph that used to look like a table.
- You can point back at the page and confirm each value.
Practical rule: extract is finished when you have data you can use. A searchable PDF is a different result. A chat answer is a different result.
Extract vs OCR vs chat
Important distinction: OCR, extract, and chat start from the same PDF and do three different jobs. Mixing them up wastes a step.
OCR (text layer)
OCR adds a hidden text layer so a scanned page becomes searchable and selectable. You still have a PDF. You can use Cmd-F / Ctrl+F. You have not pulled fields into a record.
If the file is a picture of paper, do that first. See OCR PDF: Make a Scanned File Searchable. Extract cannot invent reliable fields from pixels that were never turned into text.
Extract (fields and tables)
Extract pulls specific data points out of the document. On PDF.ai, field extraction is AI-powered extraction with custom prompts: names, dates, amounts, and other fields you define. The live extract page is built for that. It returns structured output (JSON) plus citations that link a field back to the source.
Parse, on the live parser page, turns the file into structured JSON with headings, paragraphs, tables, and figures. Use that when the thing you need is layout and tables as data, not a single named field.
Chat (answers)
Chat answers a question about the file. That is useful when you want an explanation. It is the wrong shape when you need the same fields from every invoice this week. A sentence is not a schema.
Job | What you get | What you still do not have |
OCR | A searchable, selectable PDF | Named fields or table rows |
Extract | Named fields, and tables as rows | A rewritten document, or a chat answer |
Chat | An answer in sentences | A schema you can drop into a sheet |
Practical rule: if the next step is a spreadsheet, a database, or an API payload, you are in extract. If the next step is Cmd-F, you are in OCR. If the next step is a question, you are in chat.
Decide what you need before you upload
Do not start by asking a tool to “get everything out of this PDF.” Name the output.
- List the fields you will actually use. Invoice number, date, total is a list. “All the text” is not.
- Decide whether you also need tables. Line items are a table. A paragraph of terms is not.
- Open the file and try to select a word. If nothing highlights, you have a scan. OCR first, then extract.
A digital PDF already has text. You can often copy one value by hand. Extract is worth it when you need the same fields every time, or when the table will not paste cleanly.
Practical rule: write the field names down before you upload. The extract step should fill that list, not invent a new one.
How to extract data from a PDF in practice
You can do this in a browser. You do not need to install a desktop app or write a script for a single file.
1. Confirm the file can be read as text
Use the same 10-second test as OCR:
- Open the PDF in any viewer.
- Try to drag-select a word in the body.
- Search for a word you can clearly see.
If select and search fail, treat it as a scan. On the current PDF.ai pricing page, OCR is capped by plan: Hobby 2 pages/file, Pro 10 pages/file, Ultimate 50 pages/file, Enterprise 100 pages/file. Max upload size is 10MB (Hobby), 50MB (Pro and Ultimate), and 100MB (Enterprise). Hobby is $0, listed as free forever.
Those OCR limits are for the text-layer step. This page does not invent a separate extract page cap. Check pricing for the plan you are on before you upload a long file.
If the scan is longer than the OCR page cap, split it first or use a plan that covers the page count. Then come back to extract.
2. Extract the fields you named
Go to https://pdf.ai/extract-pdf. The page is for pulling structured data such as names, dates, amounts, and custom fields. The upload area is labeled Try your file.

Say what you want out. The product copy calls this custom prompts: you define the fields, then the extract step fills them. A typical invoice list is invoice number, date, total amount, vendor name. That matches the example schema on the extract page. Do not add fields you will not use.
If what you need is the table itself (rows of line items, not one total), use the parser instead. That page says it parses PDFs into structured data and returns JSON with headings, paragraphs, tables, and figures. It also says you can try a file with no signup.
Wait until processing finishes. Keep the original PDF until you have checked the values.
Happy path: name the fields → upload → get structured data → check against the page.
3. Check the output before you trust it
Extract can return a clean JSON object and still be wrong on one number. Check the values, not the fact that a result appeared.
- Read every field you named against the source page.
- For a table, check the first row, a middle row, and the last row you care about.
- Check tokens that are easy to mangle: names, dates, amounts, invoice numbers (0/O, 1/l, 5/S).
- If the extract response includes citations, use them to jump back to the source segment. The extract page shows citations with a page number and a link into the schema.
Practical rule: a filled field is not a verified field. Read the line on the PDF before you paste the value into a sheet.
If a field is empty, the name you used may not match what the document calls it, the value may sit in an image or stamp, or the page may still be a scan. Fix the input. Do not re-run the same file and hope.
When extract goes wrong
Most failures are the wrong job, not a mysterious model error.
- You wanted search. The file is a scan and Cmd-F fails. That is OCR. Go to OCR PDF: Make a Scanned File Searchable.
- You wanted an explanation. “What is this contract about?” is chat, not extract.
- The table is a picture. A photographed spreadsheet inside a PDF is an image. OCR the file first, then parse the table.
- The layout is a mess. Nested tables, handwriting in the amount box, stamps across a total. If you cannot read it, do not expect a clean field.
- You asked for everything. A long field list on a short letter will leave empty boxes. Cut the list to the fields you will use.
This guide will not quote an accuracy percentage. Quality depends on the page in front of you. The check is always the same: does this value match the document?
Extract vs copy-paste vs converting the file
These three get mixed up because they all start with “I need the data out of this PDF.”
Copy and paste when you need one or two values from a digital PDF and the text selects cleanly. Fast for a single invoice. Fragile for tables, and it does not scale to a folder of files.
Convert the whole file (to Excel, to Word, to a giant text dump) when you want an editable document. That is a format change. You still have to hunt for the fields.
Extract when you want named fields or table rows you can use. The PDF can stay a PDF. The output is the data.
A simple split:
- Find the total on this invoice once. → copy it, if the text selects.
- Get invoice number, date, total, vendor from every PDF this week. → extract.
- Get the line-item table into rows. → parse the table.
- Make this scan searchable. → OCR, then extract if you still need fields.
Practical rule: if the output you want is a small set of fields or a table, stop at extract. Do not turn a data problem into a full-document conversion.
A short action plan
- Write down the fields (and tables) you actually need.
- Open the file and try select plus Cmd-F / Ctrl+F.
- If nothing highlights, OCR first. Check page count and file size on pricing.
- Upload at https://pdf.ai/extract-pdf for named fields, or https://pdf.ai/parser when you need tables and layout as JSON.
- Read every returned value against the source page.
- Keep the original PDF until the values you will use are correct.
If you want the rest of the PDF workflow after the data is out, start at PDF.ai. Extract is the structured-data step. OCR comes before it when the file is a scan. Chat is a different job.