# AI-Ready Webpage Extractor (`s3nafps/ai-ready-webpage-extractor`) Actor

Convert public web pages into clean Markdown, text, links, tables, images, metadata, and JSON-LD for AI workflows.

- **URL**: https://apify.com/s3nafps/ai-ready-webpage-extractor.md
- **Developed by:** [mohamed senator](https://apify.com/s3nafps) (community)
- **Categories:** AI, Developer tools
- **Stats:** 4 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $20.00 / 1,000 page extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## AI-Ready Webpage Extractor

Convert public web pages into clean Markdown, text, links, tables, images, metadata, and schema.org JSON-LD for AI workflows.

### What this Actor does

AI-Ready Webpage Extractor fetches public web pages, extracts the main content area, and saves structured content records to the Apify Dataset.

It is useful when you need clean page content for LLM prompts, RAG ingestion, search indexes, or content analysis.

### Use cases

- Prepare public pages for LLM prompts.
- Build RAG ingestion pipelines.
- Extract docs, blog, pricing, and marketing content.
- Collect links, headings, tables, images, metadata, and JSON-LD from public pages.

### Who this is for

AI app builders, RAG pipeline builders, researchers, data teams, documentation teams, and agencies that need clean public page content.

### Input

- `urls` - Public pages to extract. Required, 1-100 items.
- `outputFormat` - Output content format: `markdown`, `text`, or `both`. Default: `markdown`.
- `includeLinks` - Return extracted links with anchor text. Default: `true`.
- `includeTables` - Return HTML tables as structured row arrays. Default: `true`.
- `includeMetadata` - Return metadata and schema.org JSON-LD when present. Default: `true`.
- `maxPages` - Maximum pages to extract in one run. Default: `100`, maximum: `100`.

### Example input

```json
{
  "urls": ["https://docs.apify.com/platform"],
  "outputFormat": "markdown",
  "includeLinks": true,
  "includeTables": true,
  "includeMetadata": true,
  "maxPages": 10
}
```

### Output

The Actor saves one Dataset item per page.

- `sourceUrl` - Final fetched page URL.
- `finalUrl` - Same as `sourceUrl` in this version.
- `title` - Page title.
- `description` - Meta description or Open Graph description.
- `author` - Meta author when present.
- `publishedDate`, `modifiedDate` - Article or time metadata when present.
- `mainContentMarkdown` - Markdown content unless `outputFormat` is `text`.
- `mainContentText` - Plain text content when `outputFormat` is `text` or `both`.
- `headings` - Heading levels and text.
- `links` - Extracted links, or an empty array when disabled.
- `tables` - Extracted table row arrays, or an empty array when disabled.
- `images` - Image URLs and alt text from the main content.
- `schemaOrgData` - Parsed JSON-LD objects, or an empty array when disabled.
- `wordCount` - Approximate word count.
- `estimatedTokens` - Rough token estimate based on text length.
- `language` - HTML language attribute when present.
- `scrapedAt` - ISO timestamp for the run.
- `extractionStatus` - `succeeded` or `failed`.
- `errorMessage` - Failure reason when extraction fails.

### Example output

```json
{
  "sourceUrl": "https://docs.apify.com/platform",
  "finalUrl": "https://docs.apify.com/platform",
  "title": "Apify platform",
  "description": "Apify platform documentation",
  "author": null,
  "publishedDate": null,
  "modifiedDate": null,
  "mainContentMarkdown": "# Platform\n\nApify platform content...",
  "mainContentText": null,
  "headings": [{ "level": 1, "text": "Platform" }],
  "links": [{ "text": "Actors", "url": "https://docs.apify.com/platform/actors" }],
  "tables": [],
  "images": [],
  "schemaOrgData": [],
  "wordCount": 900,
  "estimatedTokens": 1200,
  "language": "en",
  "scrapedAt": "2026-07-04T00:00:00.000Z",
  "extractionStatus": "succeeded",
  "errorMessage": null
}
```

### How it works

- Reads the Actor input.
- Normalizes and deduplicates public URLs.
- Fetches each page as public HTML.
- Selects `main`, `article`, or `body` as the content root.
- Extracts Markdown, text, headings, links, tables, images, metadata, and JSON-LD.
- Saves one structured record per page to the Apify Dataset.

### Limitations

- Public pages only.
- No login, private data, cookies, CAPTCHA bypass, or restricted content.
- Static HTML extraction is used first.
- JavaScript-only sites may return limited or failed content.
- Main-content detection is heuristic and depends on page structure.
- Token estimates are approximate.

### Error handling

Invalid or empty `urls` input fails run validation. Request failures return a failed Dataset item with `extractionStatus: "failed"` and `errorMessage`. Pages with too little extracted main content are marked as failed with `Not enough main content extracted`.

### Troubleshooting

- If extracted content is too short, the page may rely on JavaScript or hide the main article content from the initial HTML.
- If you only need plain text, set `outputFormat` to `text`.
- If the Dataset is larger than expected, disable links, tables, or metadata fields you do not need.
- If metadata is missing, check whether the source page exposes meta tags or JSON-LD in public HTML.

### API usage

```bash
curl -X POST "https://api.apify.com/v2/acts/USERNAME~ai-ready-webpage-extractor/runs?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "urls": ["https://docs.apify.com/platform"],
    "outputFormat": "both",
    "includeLinks": true,
    "includeTables": true,
    "includeMetadata": true,
    "maxPages": 10
  }'
```

### Integration ideas

- Send Markdown into an LLM prompt or RAG pipeline.
- Export text records to a search index.
- Use links as crawl targets for follow-up runs.
- Store JSON-LD metadata for content enrichment.

### SEO keywords

AI webpage extractor, Markdown extractor, RAG scraper, clean text API, webpage to markdown, content extraction, docs extractor, blog extractor, Apify AI actor, public web data

### Ethical use

This Actor is designed for public data and user-provided URLs only. Do not use it to access private, login-protected, or restricted content.

# Actor input Schema

## `urls` (type: `array`):

Public pages to extract.

## `outputFormat` (type: `string`):

Choose whether to return Markdown, plain text, or both.

## `includeLinks` (type: `boolean`):

Return extracted links with anchor text.

## `includeTables` (type: `boolean`):

Return HTML tables as structured row arrays.

## `includeMetadata` (type: `boolean`):

Return metadata and schema.org JSON-LD where present.

## `maxPages` (type: `integer`):

Maximum public pages to extract in one run.

## Actor input object example

```json
{
  "urls": [
    "https://docs.apify.com/platform"
  ],
  "outputFormat": "markdown",
  "includeLinks": true,
  "includeTables": true,
  "includeMetadata": true,
  "maxPages": 100
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://docs.apify.com/platform"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("s3nafps/ai-ready-webpage-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://docs.apify.com/platform"] }

# Run the Actor and wait for it to finish
run = client.actor("s3nafps/ai-ready-webpage-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://docs.apify.com/platform"
  ]
}' |
apify call s3nafps/ai-ready-webpage-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=s3nafps/ai-ready-webpage-extractor",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/zscoWJkbwLzwV6rMc/builds/ZFvtPqbR4ZjQcCNP7/openapi.json
