# PDF to Markdown Converter: Docling Parser for AI & RAG (`raional/pdf-to-markdown-converter`) Actor

Convert PDF, DOCX, PPTX, XLSX, HTML and images to clean Markdown and structured JSON using IBM's open-source Docling library. Preserves headings, tables, and page structure. RAG-ready chunked output mode for LLM pipelines.

- **URL**: https://apify.com/raional/pdf-to-markdown-converter.md
- **Developed by:** [Raion Al](https://apify.com/raional) (community)
- **Categories:** Developer tools, Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $8.00 / 1,000 pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Extract Text from PDF: PDF to Markdown & JSON Converter for AI & RAG

**Extract text from PDF, DOCX, PPTX, XLSX, HTML, and images** into clean Markdown or structured JSON. Headings, tables, and page structure are preserved. Scanned documents are handled with OCR. Includes a **RAG-ready chunked output mode** for feeding LLM pipelines and vector stores directly.

Powered by [Docling](https://github.com/docling-project/docling), IBM's open-source document-conversion library (MIT licensed, 51,000+ GitHub stars, 2.4M+ monthly downloads).

Great for: **extract text from PDF**, **PDF text extractor**, **PDF to Markdown**, **document parser**, **extract tables from PDF**, RAG/LLM ingestion pipelines, AI knowledge bases, research paper processing, and converting legacy DOCX/PPTX archives to Markdown.

### What it does

Give it a list of document URLs, or upload a file directly: it extracts the text from each one and returns clean structured output. No files are stored beyond the run.

### Output formats

| `outputFormat` | What you get |
|---|---|
| `markdown` | Clean Markdown text |
| `json` | Full structured document (headings, tables, layout, bounding boxes) |
| `both` (default) | Markdown + JSON |
| `chunks` | RAG-ready semantic chunks, each with its heading context: ready to embed and index |

### Example input

```json
{
  "documents": [
    "https://example.com/report.pdf",
    "https://example.com/slides.pptx"
  ],
  "outputFormat": "both",
  "ocrEnabled": false
}
```

### Output

One row per document:

| Field | Description |
|---|---|
| `documentUrl` | The document converted |
| `format` | Detected format (pdf, docx, pptx, xlsx, html, image) |
| `pageCount` | Pages in the document |
| `markdown` | Markdown output (if requested) |
| `json` | Structured JSON output (if requested) |
| `chunks` / `chunkCount` | RAG chunks (if `outputFormat: "chunks"`) |
| `status` | `ok` or `error` |

### OCR

Turn on `ocrEnabled` for **scanned** PDFs and images with no embedded text layer. It's slower and billed separately. Leave it off for normal digital PDFs/DOCX/PPTX, which convert faster without it.

### Pricing

Billed **per page successfully converted** (fair for a 2-page doc vs. a 200-page one) plus a per-document OCR surcharge when OCR is used. Failed documents are **not charged**.

### FAQ

**How do I extract text from a PDF?**
Paste the PDF's URL in `documents` (or upload the file) and run. The `markdown` field in the output holds the extracted text, ready to use as-is.

**How do I convert a PDF to Markdown?**
Same run: Markdown is the default output format. The `markdown` field is clean Markdown with headings and tables preserved.

**Can it extract tables from a PDF?**
Yes. Tables are detected and preserved as Markdown tables in the `markdown` output, and as structured data in the `json` output.

**Does this work for scanned PDFs?**
Yes. Turn on `ocrEnabled`. It's slower than normal text extraction, so it's billed at a separate rate; leave it off for regular digital PDFs.

**What is the "chunks" output for?**
It splits the document into semantically coherent pieces (respecting headings and structure) sized for embedding into a vector database: the standard input shape for RAG pipelines.

**How is this different from just using Docling myself?**
Docling requires installing Python + ML dependencies and managing compute. This runs it as a hosted API: no setup, pay only for what you convert.

### Please note

Only convert documents you have the right to process. Documents are processed transiently and not retained beyond the run.

Built with [Apify Python SDK](https://docs.apify.com/sdk/python/) + [Docling](https://github.com/docling-project/docling).

# Actor input Schema

## `documents` (type: `array`):

List of document URLs to convert. Supports PDF, DOCX, PPTX, XLSX, HTML, PNG/JPG/TIFF images. Leave empty if you're only using the file upload field below.

## `uploadedDocument` (type: `string`):

Have a file on your computer instead of a URL? Upload it here. Combines with any URLs listed above.

## `outputFormat` (type: `string`):

markdown = clean Markdown text. json = full structured document (headings, tables, layout). both = markdown + json. chunks = RAG-ready semantic chunks for LLM/vector-store ingestion.

## `ocrEnabled` (type: `boolean`):

Turn on for scanned PDFs/images with no embedded text layer. Slower and priced separately. Leave off for normal digital PDFs/DOCX for faster, cheaper conversion.

## `maxDocuments` (type: `integer`):

Safety cap on how many documents to process in one run.

## Actor input object example

```json
{
  "documents": [
    "https://arxiv.org/pdf/2408.09869"
  ],
  "outputFormat": "both",
  "ocrEnabled": false,
  "maxDocuments": 50
}
```

# Actor output Schema

## `documents` (type: `string`):

One dataset row per document: documentUrl, format, pageCount, markdown, json, chunks, status.

## `summary` (type: `string`):

Counts of documents processed and errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "documents": [
        "https://arxiv.org/pdf/2408.09869"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("raional/pdf-to-markdown-converter").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "documents": ["https://arxiv.org/pdf/2408.09869"] }

# Run the Actor and wait for it to finish
run = client.actor("raional/pdf-to-markdown-converter").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "documents": [
    "https://arxiv.org/pdf/2408.09869"
  ]
}' |
apify call raional/pdf-to-markdown-converter --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=raional/pdf-to-markdown-converter",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/0LuyNYiu25jDCbrdN/builds/aaxH6m7vepx0ar7IS/openapi.json
