# Web Page & PDF to Markdown (`prizable_aster/web-page-pdf-to-markdown`) Actor

Converts public web pages and text-based PDFs into clean Markdown, plain text, and structured JSON.

- **URL**: https://apify.com/prizable\_aster/web-page-pdf-to-markdown.md
- **Developed by:** [Vaque Wei](https://apify.com/prizable_aster) (community)
- **Categories:** Developer tools, Automation, AI
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.20 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Web Page & PDF to Markdown

Convert public HTML pages and text-based PDFs into clean Markdown, plain text, and structured JSON. The Actor produces one predictable dataset item per input URL, making it easy to connect web content to AI agents, RAG pipelines, knowledge bases, and automation workflows.

### What it does

- Accepts up to 100 public HTTP/HTTPS URLs per run
- Detects HTML pages and PDFs automatically
- Extracts readable Markdown and plain text
- Follows up to five validated public redirects
- Returns stable error codes instead of failing the whole batch
- Optionally includes truncated raw HTML

### Common use cases

- Prepare documentation pages for an AI knowledge base
- Normalize public PDF handbooks for search and summarization
- Convert articles into Markdown for content workflows
- Feed clean text into RAG, MCP, Make, Zapier, or custom APIs

### Input

```json
{
  "urls": [
    "https://docs.apify.com/sdk/python/docs/quick-start",
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ],
  "maxItems": 10,
  "requestTimeoutSecs": 45,
  "maxOutputChars": 40000,
  "maxDownloadMbytes": 10,
  "includeRawHtml": false
}
```

### Output

Each URL produces one default dataset item:

```json
{
  "url": "https://example.com/article",
  "status": "ok",
  "sourceType": "html",
  "finalUrl": "https://example.com/article",
  "httpStatus": 200,
  "contentType": "text/html; charset=utf-8",
  "bytesDownloaded": 12345,
  "title": "Example article",
  "markdown": "# Example article\n\nClean content...",
  "text": "Example article\n\nClean content..."
}
```

Failed URLs still produce a normalized item with `status: "error"`, a stable `errorCode`, and a readable `error` message. Other URLs in the batch continue processing.

### Safety and limitations

- Only public HTTP and HTTPS URLs are accepted
- Private, loopback, link-local, reserved, and credential-bearing URLs are blocked
- Redirect destinations are validated before they are requested
- Downloads are limited to 10 MB by default and 25 MB maximum
- Scanned/image-only PDFs require OCR and are not supported in this version
- The Actor does not log in, click through pages, or bypass access controls

Only process content you are authorized to access and respect source website terms and applicable law.

### Local run

```powershell
python -m venv .venv
.\.venv\Scripts\python -m pip install -r requirements.txt
.\.venv\Scripts\python -m src
```

### Sync to Apify

Set `APIFY_TOKEN` for the current shell, then run:

```powershell
.\.venv\Scripts\python .\tools\deploy_actor.py
```

# Actor input Schema

## `urls` (type: `array`):

Public page or PDF URLs to process.

## `maxItems` (type: `integer`):

Optional cap on how many URLs to process from the input list.

## `requestTimeoutSecs` (type: `integer`):

HTTP timeout used for downloading pages or PDFs.

## `maxOutputChars` (type: `integer`):

Truncate extracted markdown/text to this many characters per item.

## `maxDownloadMbytes` (type: `integer`):

Stop downloading a URL when its response exceeds this size.

## `includeRawHtml` (type: `boolean`):

When enabled, store the fetched HTML in the output for HTML pages.

## Actor input object example

```json
{
  "urls": [
    "https://docs.apify.com/sdk/python/docs/quick-start"
  ],
  "maxItems": 10,
  "requestTimeoutSecs": 45,
  "maxOutputChars": 40000,
  "maxDownloadMbytes": 10,
  "includeRawHtml": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://docs.apify.com/sdk/python/docs/quick-start"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("prizable_aster/web-page-pdf-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://docs.apify.com/sdk/python/docs/quick-start"] }

# Run the Actor and wait for it to finish
run = client.actor("prizable_aster/web-page-pdf-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://docs.apify.com/sdk/python/docs/quick-start"
  ]
}' |
apify call prizable_aster/web-page-pdf-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=prizable_aster/web-page-pdf-to-markdown",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/5MfAQXbrZuIFFhfot/builds/Aq5mh7uxcN99UmxKW/openapi.json
