# Website Content to Markdown (LLM-ready) (`vivid_astronaut/website-content-to-markdown`) Actor

Turn any website into clean, LLM-ready Markdown for RAG pipelines, AI agents and knowledge bases. Scrape single pages or crawl entire sites. Compliance-first: robots.txt honored.

- **URL**: https://apify.com/vivid\_astronaut/website-content-to-markdown.md
- **Developed by:** [BRAINIALL Team](https://apify.com/vivid_astronaut) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 page converted to markdowns

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Website Content to Markdown (LLM-ready) — by Brainiall

Turn any website into **clean, structured Markdown** that is ready to feed into LLMs, RAG pipelines, AI agents, vector databases and knowledge bases — without writing a single selector.

Give it one or more URLs and get back distraction-free Markdown: headings, paragraphs, lists and links preserved; navigation chrome, cookie banners, scripts and boilerplate stripped out.

### What it does

- **Single-page scrape** — convert exact URLs to Markdown (default mode).
- **Site crawl** — set *Max pages per site* above 1 and each start URL becomes the entry point of a bounded crawl (up to 25 pages, configurable link depth). Every page becomes one dataset item.
- **Metadata included** — page title, language, description and word count with every item.
- **Links on demand** — optionally return the hyperlinks found on each page.

Powered by the **Brainiall Web engine** ([api.brainiall.com](https://app.brainiall.com)) — a production web-intelligence service built for AI workloads.

### Who it's for

- **RAG builders** — ingest documentation sites, blogs and product pages straight into your vector store. Markdown chunks cleanly and embeds better than raw HTML.
- **AI agent developers** — give your agents reliable, token-efficient web content instead of noisy HTML.
- **Data & content teams** — archive or migrate site content as portable Markdown.
- **LLM fine-tuning** — collect clean text corpora from public websites.

### Compliance-first by design

This Actor inherits the ethics posture of the Brainiall Web engine:

- **robots.txt is honored** — pages disallowed for crawlers are skipped, and crawl-delay directives are respected.
- **No bot-detection bypassing** — sites that choose to block automated access stay blocked.
- **SSRF-guarded** — internal/private addresses are never fetched.

If your compliance team asks how the data was collected, you have a clean answer.

### Input

```json
{
    "startUrls": [
        { "url": "https://docs.example.com" },
        { "url": "https://blog.example.com/post" }
    ],
    "maxPagesPerUrl": 10,
    "maxDepth": 2,
    "includeLinks": false
}
```

| Field | Description |
|-------|-------------|
| `startUrls` | Pages (or crawl entry points) to convert. |
| `maxPagesPerUrl` | `1` = scrape only the exact URL. `2-25` = crawl the site from that URL. |
| `maxDepth` | Link hops allowed from the start URL (crawl mode). |
| `includeLinks` | Also return hyperlinks found on the page (single-page mode). |

### Output

One dataset item per page:

```json
{
    "url": "https://example.com/",
    "requested_url": "https://example.com",
    "title": "Example Domain",
    "markdown": "# Example Domain\n\nThis domain is for use in documentation examples...",
    "word_count": 19,
    "metadata": { "description": "", "lang": "en" },
    "http_status": 200
}
```

Export as JSON, CSV, Excel or consume via the Apify API — ready for LangChain, LlamaIndex or any custom pipeline.

### Pricing

You pay **per result** (per page successfully converted to Markdown). Pages that fail or contain no extractable content are **not** charged. No subscriptions, no minimums — costs scale exactly with what you extract.

### Tips

- For documentation sites, start at the docs root with `maxPagesPerUrl: 25`, `maxDepth: 3` to capture a whole section in one run.
- Keep `includeLinks` off unless you need them — smaller items mean faster downstream processing.
- Deduplicate by `url` if you crawl overlapping start URLs.

***

Built and maintained by [Brainiall](https://www.brainiall.com) — production AI APIs for speech, documents, vision and the web.

# Actor input Schema

## `startUrls` (type: `array`):

Pages or sites to convert to Markdown. With <b>Max pages per site</b> = 1 each URL is scraped as a single page; with a higher value each URL is used as the entry point of a site crawl.

## `maxPagesPerUrl` (type: `integer`):

1 = scrape only the exact URL. 2-25 = crawl the site starting from the URL, following internal links, up to this many pages. Each page becomes one dataset item.

## `maxDepth` (type: `integer`):

How many link hops away from the start URL the crawler may go (crawl mode only).

## `includeLinks` (type: `boolean`):

Also return the hyperlinks found on each page (single-page scrape mode only).

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://example.com"
    }
  ],
  "maxPagesPerUrl": 1,
  "maxDepth": 2,
  "includeLinks": false
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://example.com"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("vivid_astronaut/website-content-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://example.com" }] }

# Run the Actor and wait for it to finish
run = client.actor("vivid_astronaut/website-content-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://example.com"
    }
  ]
}' |
apify call vivid_astronaut/website-content-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=vivid_astronaut/website-content-to-markdown",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/ET8IhlMxHRMWwMh6D/builds/AC6TEkb1TEgEI0Etv/openapi.json
