# PDF Extractor: PDF → Clean Markdown + JSON for LLM/RAG (`boxbox10/pdf-extractor`) Actor

Turn any PDF URL into clean, LLM-ready Markdown + structured JSON (title, metadata, per-page text, page count, word count, token count). Perfect for RAG pipelines, AI agents, and LLM document ingestion.

- **URL**: https://apify.com/boxbox10/pdf-extractor.md
- **Developed by:** [Marvin Eguilos](https://apify.com/boxbox10) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## PDF Extractor — PDF → Clean Markdown + JSON for LLM & RAG

**Give it a PDF URL. Get back clean, LLM-ready Markdown and structured JSON.** Title, metadata, per-page text, page count, word count, and an accurate token count — every document, one tidy result. Built for RAG pipelines, AI agents, and anyone who needs the *text* of a PDF without wrestling with binary parsing, broken encodings, or page-layout noise.

> Feed your PDFs to your LLM the way it wants to be fed: as clean Markdown, chunkable by page, with the token budget already counted.

***

### ✨ What it does

- **Robust PDF text extraction** — powered by Mozilla's `pdf.js` (via `unpdf`), the same battle-tested engine that renders PDFs in Firefox. Handles multi-page documents, embedded fonts, and complex layouts.
- **Clean Markdown** — normalized whitespace, collapsed noise, optional per-page separators (`--- Page N ---`) so you can chunk on natural document boundaries.
- **Structured JSON** — `title`, `metadata` (author, subject, keywords, creator, producer, creation/modification dates, PDF version), `pageCount`, `pages[]` (per-page text), `wordCount`, `tokenCount`, `fileSizeBytes`, `fetchedAt`.
- **Accurate token counts** — counted with the GPT/`cl100k`-family tokenizer so you know exactly how much context each PDF costs *before* you send it to a model.
- **Token budgeting** — optional `maxTokens` truncates output to fit your context window.
- **Page-level chunking** — the `pages[]` array gives you clean text per page, ideal for granular RAG retrieval and citations.
- **Robust by design** — one bad URL never kills the run. Failed downloads/parses return a clean error record (and are **never charged**).
- **Safe & efficient** — file-size guard, `%PDF-` magic-byte validation, and content-type checks reject non-PDFs before wasting compute.

***

### 🎯 Use cases

| You want to… | This Actor gives you… |
|---|---|
| **Build a RAG knowledge base from PDFs** | Clean Markdown + per-page text with token counts, ready to embed. |
| **Ingest research papers (arXiv, journals)** | Structured text + metadata your pipeline can index and cite. |
| **Feed reports / whitepapers to an LLM** | Pre-counted tokens so you never blow the context window. |
| **Give an AI agent document context** | Structured JSON your agent can reason over — no binary noise. |
| **Convert PDF docs to Markdown** | Portable, diff-friendly Markdown you can drop into a repo or wiki. |
| **Chunk documents by page for citations** | A `pages[]` array with clean text and page numbers. |

***

### 📥 Input

| Field | Type | Default | Description |
|---|---|---|---|
| `urls` | string\[] | — **(required)** | One or more direct links to PDF files. Each successful PDF is one result. |
| `outputFormat` | `both` | `markdown` | `json` | `both` | Include Markdown, the JSON fields, or both. |
| `includePageText` | boolean | `true` | Include a `pages[]` array with per-page text in the JSON output. |
| `pageSeparators` | boolean | `true` | Insert `--- Page N ---` markers between pages in the Markdown. |
| `maxPagesPerRun` | integer | `1000` | Safety cap on PDF URLs processed per run. |
| `maxTokens` | integer | `0` | Truncate Markdown to ~N tokens (`0` = no limit). |
| `maxFileSizeMb` | integer | `100` | Skip PDFs larger than this size (recorded as failures, never charged). |

#### Example input

```json
{
  "urls": [
    "https://arxiv.org/pdf/1706.03762",
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ],
  "outputFormat": "both",
  "includePageText": true,
  "pageSeparators": true,
  "maxTokens": 0
}
```

***

### 📤 Output

One dataset item per PDF. Successful example (truncated):

```json
{
  "url": "https://arxiv.org/pdf/1706.03762",
  "finalUrl": "https://arxiv.org/pdf/1706.03762",
  "statusCode": 200,
  "title": "Attention Is All You Need",
  "metadata": {
    "title": "Attention Is All You Need",
    "author": "Ashish Vaswani et al.",
    "creator": "LaTeX with hyperref",
    "producer": "pdfTeX",
    "creationDate": "D:20170612...",
    "pdfVersion": "1.5"
  },
  "pageCount": 15,
  "fileSizeBytes": 2215244,
  "wordCount": 8823,
  "tokenCount": 13140,
  "fetchedAt": "2026-07-19T02:17:19.954Z",
  "markdown": "Attention Is All You Need\n\n...\n\n---\n\n**Page 2**\n\n...",
  "pages": [
    { "page": 1, "text": "Attention Is All You Need\n\n..." }
  ]
}
```

Failed URL (returned to a separate `failures` dataset, **not charged**):

```json
{
  "url": "https://not-a-real-domain-xyz.com/file.pdf",
  "finalUrl": "https://not-a-real-domain-xyz.com/file.pdf",
  "statusCode": null,
  "error": "getaddrinfo ENOTFOUND not-a-real-domain-xyz.com",
  "fetchedAt": "2026-07-19T02:17:19.831Z"
}
```

***

### 🔌 Call it from code (Apify API)

```bash
curl -X POST "https://api.apify.com/v2/acts/YOUR_USERNAME~pdf-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://arxiv.org/pdf/1706.03762"]}'
```

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const { defaultDatasetId } = await client
    .actor('YOUR_USERNAME/pdf-extractor')
    .call({ urls: ['https://arxiv.org/pdf/1706.03762'] });
const { items } = await client.dataset(defaultDatasetId).listItems();
console.log(items[0].markdown);
```

***

### 💸 Pricing (Pay-Per-Event)

| Event | Price |
|---|---|
| Actor start | **$0.05** per run |
| Extracted PDF | **$0.003** per successfully extracted PDF |

- **You only pay for PDFs that succeed** — failed URLs are never charged.
- **🎁 Free tier:** free-plan users' platform usage is covered by Apify, so you can try it and run small jobs at **no cost** before scaling up.
- Structured per-page JSON + accurate token counts, purpose-built for LLM/RAG ingestion — at a low, predictable price point.

***

### ⚖️ Acceptable use

This is a general-purpose format-conversion tool: **you supply the PDF URLs and are responsible for having the right to download and use the content you submit.** The Actor identifies itself with a descriptive User-Agent and only extracts text from documents you point it at. It does not target any single platform's private API and does not harvest personal data as a feature. Please comply with each source's terms of service and applicable law.

***

### 🧱 Under the hood

Node.js · [unpdf](https://github.com/unjs/unpdf) (Mozilla `pdf.js` core, serverless-friendly) · [gpt-tokenizer](https://github.com/niieani/gpt-tokenizer) · Apify SDK. Pure-JS extraction — no headless browser, no system libraries. Stateless — nothing is stored between runs.

# Actor input Schema

## `urls` (type: `array`):

One or more direct links to PDF files. Each successfully extracted PDF is one billable result.

## `outputFormat` (type: `string`):

What to include per result: Markdown only, JSON fields only, or both.

## `includePageText` (type: `boolean`):

Include a structured `pages[]` array with the raw text of each page in the JSON output. Useful for page-level chunking in RAG.

## `pageSeparators` (type: `boolean`):

Insert a horizontal-rule page marker (e.g. '--- Page 2 ---') between pages in the Markdown output.

## `maxPagesPerRun` (type: `integer`):

Safety cap on how many PDF URLs are processed in a single run.

## `maxTokens` (type: `integer`):

Optional. Truncate the Markdown output to approximately this many tokens (0 = no limit). Useful for fitting LLM context windows.

## `maxFileSizeMb` (type: `integer`):

Skip PDFs larger than this size (in megabytes) to protect compute. Oversized files are recorded as failures and never charged.

## Actor input object example

```json
{
  "urls": [
    "https://arxiv.org/pdf/1706.03762"
  ],
  "outputFormat": "both",
  "includePageText": true,
  "pageSeparators": true,
  "maxPagesPerRun": 1000,
  "maxTokens": 0,
  "maxFileSizeMb": 100
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
        "https://arxiv.org/pdf/1706.03762"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("boxbox10/pdf-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
        "https://arxiv.org/pdf/1706.03762",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("boxbox10/pdf-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
    "https://arxiv.org/pdf/1706.03762"
  ]
}' |
apify call boxbox10/pdf-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=boxbox10/pdf-extractor",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/1ZMLm2PMkja2dlKYw/builds/smDCjQhX0hUciwhD6/openapi.json
