# PDF & DOCX to Markdown — Document Extractor for LLM/RAG (`fetchbase/document-to-markdown`) Actor

Convert PDF and Word (DOCX) documents into clean Markdown, text, or JSON. Smart PDF paragraph reflow, page markers for RAG citations, full DOCX structure (headings, lists, tables), custom auth headers. No browser — parses in seconds. Charged per page processed — no startup fee.

- **URL**: https://apify.com/fetchbase/document-to-markdown.md
- **Developed by:** [Fetchbase](https://apify.com/fetchbase) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 document page processeds

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## PDF & DOCX to Markdown — Document Text Extractor for LLMs

Convert **PDF and Word (DOCX) documents** into clean **Markdown, plain text, or
JSON** — ready for RAG pipelines, LLM prompts, embeddings, and datasets. Point
it at your document URLs and get back structured, chunk-ready content.

**Billing is per page actually processed** (about ⅓ the price of hosted parsing
APIs). No startup fee. Failed documents cost $0.

### Quick start

```json
{
  "fileUrls": ["https://example.com/whitepaper.pdf"]
}
```

→ one dataset item per document:

```json
{
  "url": "https://example.com/whitepaper.pdf",
  "fileType": "pdf",
  "pages": 15,
  "wordCount": 6112,
  "title": "Attention Is All You Need",
  "markdown": "The dominant sequence transduction models are based on…"
}
```

### What it does well

| | |
|---|---|
| **PDF** | Page-aware text extraction with **smart paragraph reflow** (fixes the broken-mid-sentence line breaks raw PDF extractors produce); optional `--- Page N ---` markers so your chunks can cite pages; document title from PDF metadata |
| **DOCX** | Full structure preserved — headings, lists, **tables** (GFM), links, bold/italic — via real .docx parsing, not text-dump |
| **Outputs** | Markdown / plain text / JSON (JSON includes per-page text array for PDFs) |
| **Protected files** | Custom request headers (e.g. `Authorization: Bearer …`) for presigned or authenticated URLs you have access to |
| **Cost control** | `maxPagesPerDoc` cap (default 500) so a giant PDF can't surprise you |
| **Speed** | No browser involved — documents parse in seconds on the lightweight runtime |

### Built for AI pipelines

- **RAG ingestion:** fetch → `markdown` → chunk → embed. Use `pageMarkers` to keep
  page-level citations.
- **Whole knowledge bases:** pass hundreds of URLs in one run; each becomes one
  clean dataset record.
- **Agents:** call via the Apify API or MCP server to give an agent readable
  documents instead of binary blobs.
- Pairs with **[Website to Markdown](https://apify.com/fetchbase/website-to-markdown)**
  from the same publisher — web pages + documents, one uniform Markdown output
  for your whole ingestion pipeline.

```bash
curl -X POST "https://api.apify.com/v2/acts/fetchbase~document-to-markdown/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"fileUrls": ["https://example.com/report.pdf"], "outputFormat": "json"}'
```

### Pricing

A small fixed price **per page processed** (PDF pages; DOCX bills as ~500-word
page equivalents, minimum 1), plus standard Apify platform usage — which is tiny
here, since no browser is involved. Failed documents are never billed.

### Current limits (roadmap)

- **Scanned/image-only PDFs are not OCR'd yet** — they return a clear error and
  are not billed. OCR support is planned; ask for it in the issues tab if you
  need it and it moves up the list.
- Complex multi-column PDF layouts extract in reading order on a best-effort
  basis (this is a hard limit of the PDF format itself).
- DOC (legacy Word), ODT and image formats are not yet supported.

### FAQ

**Where do my documents go?** They're fetched, parsed in the run's container,
and the extracted content is stored only in your own run's dataset. Nothing is
retained by the actor. Only submit documents you have the right to process.

**A document failed — was I charged?** No. Only successfully processed pages
are billed.

***

Need OCR, ODT, or another format? **Open an issue** — this actor iterates fast.
If it saves you time, a ⭐ review helps others find it.

### More tools by Fetchbase

Part of a suite of fast, no-nonsense web utilities — all pay-per-result, charged only on success, no startup fee:

- [Website Screenshot API](https://apify.com/fetchbase/web-screenshot-pro) — full-page PNG / JPEG / WebP / PDF
- [Website to Markdown](https://apify.com/fetchbase/website-to-markdown) — clean Markdown for LLMs & RAG
- [PDF & DOCX to Markdown](https://apify.com/fetchbase/document-to-markdown) — documents → Markdown for RAG
- [SEO Audit + Core Web Vitals](https://apify.com/fetchbase/website-seo-audit) — scored on-page audit with fixes
- [Website Performance Audit](https://apify.com/fetchbase/website-performance-audit) — bulk Core Web Vitals & page speed
- [Tech Stack Detector](https://apify.com/fetchbase/tech-stack-detector) — CMS, frameworks, analytics, hosting
- [Domain, DNS & WHOIS Lookup](https://apify.com/fetchbase/domain-dns-intelligence) — records, registration, SSL
- [RSS Feed Reader](https://apify.com/fetchbase/rss-feed-reader) — RSS / Atom / JSON → normalized JSON
- [Job Postings API](https://apify.com/fetchbase/job-postings-scraper) — Greenhouse, Lever, Ashby & more

### Use with AI agents (MCP)

This Actor is callable by AI agents through the [Apify MCP server](https://mcp.apify.com/). Agents in Claude, Cursor, Windsurf, LangGraph, CrewAI and others can discover it via `search-actors` and run it as a tool — its inputs and outputs are fully described in the schema for reliable agent use.

***

#### Was this Actor useful?

If it saved you time, please consider leaving an honest review. Reviews are the main way
independent Actors get discovered on Apify — a single one makes a real difference, and
critical feedback is just as welcome as praise.

Something missing or broken? Open an issue instead and it will get fixed.

# Actor input Schema

## `fileUrls` (type: `array`):

Direct links to PDF or DOCX files (public URLs, presigned S3/GCS links, or Apify key-value-store record URLs). Each document produces one dataset item; billing is per page processed.

## `outputFormat` (type: `string`):

markdown (best for LLMs/RAG), text (plain), or json (markdown + text + per-page breakdown + metadata).

## `pageMarkers` (type: `boolean`):

Insert '--- Page N ---' separators between PDF pages so chunkers and citations can reference page numbers.

## `maxPagesPerDoc` (type: `integer`):

Safety cap: stop processing a PDF after this many pages (protects against huge documents and caps your cost).

## `requestHeaders` (type: `object`):

Optional headers sent when downloading the documents, e.g. {"Authorization": "Bearer …"} for protected files you have access to.

## `timeoutSecs` (type: `integer`):

Max time to download each document.

## Actor input object example

```json
{
  "fileUrls": [
    "https://example.com/report.pdf",
    "https://example.com/contract.docx"
  ],
  "outputFormat": "markdown",
  "pageMarkers": false,
  "maxPagesPerDoc": 500,
  "timeoutSecs": 60
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "fileUrls": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("fetchbase/document-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "fileUrls": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"] }

# Run the Actor and wait for it to finish
run = client.actor("fetchbase/document-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "fileUrls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ]
}' |
apify call fetchbase/document-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=fetchbase/document-to-markdown",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/xYCmHgHc2JEPmWJJg/builds/L7uxmnFaCQMwpDm98/openapi.json
