PDF OCR API: Run OCR from Your Code

Publish date
Sep 4, 2026
AI summary
The page explains when to use a browser OCR tool versus the PDF OCR API, outlines plan limits and API credit costs, provides step‑by‑step instructions for obtaining an API key and calling the Parse API with standard quality, and gives practical tips for verifying OCR output and choosing the right tool for one‑off searchable PDFs or automated code‑based OCR workflows.
Language
People type pdf ocr api when a browser upload is not enough. They need OCR inside an app, a batch job, or a pipeline: send a scanned PDF, get text (or structured content) back, then keep going in code.
This page is that job. The broader map of scanned vs digital PDFs is OCR PDF: Make a Scanned File Searchable. The same OCR job without the API focus is How to OCR a PDF or OCR a Scanned PDF. For plan caps on the free browser tool, see Free OCR PDF.

What a PDF OCR API is for

OCR reads the picture of each page and writes text the computer can use. An OCR API does that over HTTP so your code can call it.
After a clean OCR-oriented parse:
  • You get text you can search, index, or pass to the next step
  • Scanned pages stop being “image only”
  • You are not yet extracting named invoice fields (that is extract), and you are not chatting with the file
On PDF.ai, OCR for developers is exposed through the Parse API with quality set to standard. The docs label that path Standard (OCR). There is not a separate public endpoint named only /ocr. Parse returns markdown plus structured contents (and related fields). That is the OCR-in-code path this guide uses.
Practical rule: if the next step is Cmd-F in a downloaded file, the browser OCR tool may be enough. If the next step is your server, use the API.

Browser OCR vs PDF OCR API

Path
Best when
What you get
Browser tool
One file, no install, no key
A searchable PDF you download
Parse API (standard)
Automation, batches, your product
JSON / markdown from POST /api/v2/parse
The free browser tool lives at https://pdf.ai/tools/ocr-pdf. No signup for that path. It accepts PDF, PNG, JPG, and WebP.
The API path uses an API key and credits. Docs: Parse and credit usage. You can also try the parse playground at https://pdf.ai/parser or https://pdf.ai/parse-pdf.

Limits you should check first

Two different limit tables exist. Do not mix them up.

Web app OCR (pricing page)

On PDF.ai pricing, OCR is capped by plan for the product:
  • Hobby: Free forever ($0), OCR 2 pages/file, max file size 10MB
  • Pro: OCR 10 pages/file, max file size 50MB
  • Ultimate: OCR 50 pages/file, max file size 50MB
  • Enterprise: OCR 100 pages/file, max file size 100MB
Those caps apply to the product OCR column on pricing. They are not the API credit table.

API credits (developer docs)

From the live credit usage and Parse docs:
  • Parse Standard (OCR): 1 credit/page
  • Parse Advanced (VLM): 2 credits/page
  • Extra credits may apply if images need LLM analysis (llm enabled)
  • Cached parse results with the same settings use 0 credits
  • Free API plan: $0/month for 200 credits/month, no credit card required
  • Paid API examples on that page: $49 / 3,000, $99 / 10,000, $249 / 30,000, $599 / 100,000 credits per month
Generate the API key from the developer page. The docs say the key is shown only once when generated.
Practical rule: pick the browser tool when you need a searchable file once. Pick the API when OCR must run without a human at the upload box.

How to call a PDF OCR API (Parse, standard)

1. Get an API key

Create a key on the developer page. Store it as a secret. Send it as the X-API-Key header.

2. POST to Parse with standard quality

Endpoint from the docs: POST https://pdf.ai/api/v2/parse
Content type: multipart/form-data.
Useful parameters (from Parse docs):
  • file or url (or docId when you already have a cached document id)
  • quality: use standard for the OCR path (default). advanced uses VLM instead
  • lang_list: languages for OCR when using standard (default ["en"])
  • llm: optional image LLM processing (default false)
Example shape (Python), matching the docs sample:
import requests

url = "https://pdf.ai/api/v2/parse"
headers = {"X-API-Key": "YOUR_API_KEY"}

with open("/path/to/document.pdf", "rb") as f:
    files = {"file": f}
    data = {
        "quality": "standard",
        "lang_list": '["en"]'
    }
    response = requests.post(url, headers=headers, files=files, data=data)

print(response.json())
A successful response includes fields such as success, markdown, contents, images, pageCount, and docId (see the Parse docs for the full schema). Keep the original PDF until you have checked the text you care about.

3. Check the output before you trust it

OCR can return text and still be wrong on one character.
  • Search for a word from the first page and a later page you care about
  • Check tokens that are easy to mangle: names, dates, amounts, IDs (0/O, 1/l, 5/S)
  • If the page was a bad scan, fix the scan (or split a file that is too long for your path) before you blame the API
Practical rule: a non-empty markdown string is not a verified page. Read the lines you will use.

When the API path is the wrong tool

  • You need layout as structured blocks for RAG. Still Parse, but you are using the full parse result, not “make this file searchable” alone. The playground is https://pdf.ai/parser.
  • You wanted a generic “PDF API” overview. That is a different topic. This page stays on OCR via API.

A short action plan

  1. Decide: one-off searchable file, or OCR inside your code.
  1. For one-off: open https://pdf.ai/tools/ocr-pdf, check pricing page/size caps, download, then prove search.
  1. For code: create an API key, call POST https://pdf.ai/api/v2/parse with quality=standard, check credit usage.
  1. Verify text from early and late pages before you wire the next step.
  1. Keep the original PDF until the text you will use is correct.
If you want the file in the rest of your PDF workflow after the text is available, start at PDF.ai. OCR first when the file is a scan.