Adobe PDF Extract API: JSON and Markdown

Unlock the structure and content elements of any PDF with a web service powered by Adobe Sensei's machine learning.

Key features of Adobe PDF Extract API

EMPTY_ALT

Flexible JSON and Markdown outputs

Extract text, tables, and figures from PDFs using two output options powered by the same Adobe extraction technology: detailed structured JSON from the Extract PDF endpoint or clean, LLM-friendly Markdown from the PDF to Markdown endpoint.

EMPTY_ALT

Document structure understanding

Classify text objects such as headings, lists, footnotes, and paragraphs that may span multiple columns or pages. Capture text fonts and styles, positioning, and the natural reading order of all objects.

EMPTY_ALT

Highly accurate results

Adobe Sensei AI technology delivers highly accurate data extraction across a broad range of document types – both native and scanned PDFs – without requiring custom ML templates or model training.

EMPTY_ALT

Platform agnostic

Adobe’s PDF Extract API is RESTful and can be used to seamlessly integrate with any cloud platform or on-premise application.

EMPTY_ALT

See structured JSON extraction in action.

Check out the interactive demo that shows a sample PDF input and the JSON output side-by-side. Click on a section of the PDF to see the corressponding JSON output. You can extract a variety of elements such as paragraphs, headers, tables, and figures/images.

Turn your PDF into rich data.

PDF Extract provides two output formats through separate endpoints, both powered by the same underlying Adobe extraction technology:

  • Markdown, PDF to Markdown endpoint (/operation/pdftomarkdown): Returns well-formatted, LLM-friendly Markdown that preserves document structure and reading order. Tables are converted to Markdown syntax, and figures can be included as base64-embedded images.
  • Structured JSON, Extract PDF endpoint (/operation/extractpdf): Returns detailed content and document structure data in JSON. Tables can also be output as CSV or XLSX files, and figures as PNG files.

Get the document structure, not just the characters.

Adobe PDF Extract API is powered by Adobe Sensei, an industry-leading Artificial Intelligence (AI) and Machine Learning (ML) network. This enables a rich understanding of document structure, including the identification of elements, position, connections relative to other elements, and the reading order.

Get started in minutes

Start with the Free Tier and get 500 free Document Transactions per month.

Step 1

Obtain free credentials

Explore other Adobe Acrobat Services APIs

EMPTY_ALT

Services
Create a PDF from Microsoft Office documents, protect the content, and export to other formats.

EMPTY_ALT

Generate
Generate PDF and Word documents from custom Word templates.

EMPTY_ALT

Embed
Embed high-fidelity PDFs in web apps with analytics.

We're ready to help

Have questions about the Acrobat Services APIs?

  • Privacy
  • Terms of Use
  • Do not sell or share my personal information
  • AdChoices
Copyright © 2026 Adobe. All rights reserved.