# LLM-Ready Web Scraper – RAG & Vertical Data Extraction (`conceivable_extension/llm-ready-web-scraper`) Actor

Scrapes any URL and returns clean LLM-ready content. Strips ads, nav, and boilerplate. Returns markdown, chunked text, token estimates, and metadata. Vertical modes for Legal, Medical, Property, E-commerce, Research, and News. Firecrawl alternative at $0.005 per URL.

- **URL**: https://apify.com/conceivable\_extension/llm-ready-web-scraper.md
- **Developed by:** [joseph fadero](https://apify.com/conceivable_extension) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 1 total users, 0 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 url crawleds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## LLM-Ready Web Scraper – RAG Data Extraction with Vertical Processing

**The affordable Firecrawl alternative. $0.005 per URL. No subscription.**

Scrapes any public URL and returns clean, structured content optimised for LLMs and RAG pipelines — stripped of navigation, ads, cookie banners, and HTML boilerplate.

### What makes it different

- **Vertical processing modes** — Legal, Medical, Property, E-commerce, Research, and News modes apply domain-specific extraction rules for better content quality
- **RAG-ready chunking** — splits content into configurable token-sized chunks ready for embedding
- **Token estimation** — every result includes estimated token count so you know your LLM context usage upfront
- **Pay per URL** — $0.005/URL, no subscription

### Use cases

- Feed RAG pipelines with fresh web content for Claude, GPT-4, or LlamaIndex
- Build AI agents that need live web data
- n8n/Make: scrape URLs from a spreadsheet → get clean markdown → send to your LLM
- Research aggregation: scrape multiple sources → chunk → embed → search
- Legal research: extract clean text from case law and statutes
- Property analysis: extract listing descriptions for AI comparison

### Pricing

| Event | Price |
|---|---|
| Run started | $0.05 |
| URL crawled (no chunks) | $0.005 |
| URL crawled (with chunking) | $0.008 |
| URL failed | $0.001 |

**100 URLs = $0.55 total. Firecrawl Hobby plan: $19/month for 500 URLs.**

### Input

| Field | Default | Description |
|---|---|---|
| urls | required | Array of URLs to scrape |
| outputFormat | markdown | markdown / plaintext / json |
| vertical | general | general / legal / medical / property / ecommerce / research / news |
| chunkContent | false | Split into RAG-sized chunks |
| chunkTokenSize | 512 | Target tokens per chunk (128–4096) |
| includeMetadata | true | Include title, author, dates, word/token count |
| removeElements | \[] | Extra CSS selectors to strip |
| followLinks | false | Follow internal links from starting URLs |
| maxDepth | 1 | Link follow depth (1–3) |
| maxPagesPerUrl | 10 | Max pages per starting URL |

### Output fields

- `url`, `sourceUrl`, `crawledAt`
- `title`, `description`, `author`, `publishDate`, `language`
- `wordCount`, `estimatedTokens`
- `content` — clean text in chosen format
- `vertical` — which extraction mode was applied
- `chunks` — array of `{ index, content, tokenEstimate }` when chunking enabled
- `status` — success / failed / partial
- `chargedEvent`

### Example n8n workflow

Apify node → this actor → Claude AI node → Google Sheets

# Actor input Schema

## `urls` (type: `array`):

List of URLs to scrape and convert to LLM-ready content. Supports any public web page.

## `outputFormat` (type: `string`):

Format for the extracted content. Markdown is recommended for LLMs.

## `vertical` (type: `string`):

Optimises extraction rules for the content type. General works for all pages.

## `chunkContent` (type: `boolean`):

Split content into chunks suitable for embedding. Set chunk size with chunkTokenSize.

## `chunkTokenSize` (type: `integer`):

Approximate token size per chunk when chunking is enabled. 512 is standard for most embedding models.

## `includeMetadata` (type: `boolean`):

Include title, description, author, publish date, word count, and estimated token count in output.

## `removeElements` (type: `array`):

CSS selectors for elements to strip from the page before extraction. Example: '.sidebar', '#comments'

## `followLinks` (type: `boolean`):

If enabled, follows internal links from each starting URL and scrapes linked pages too. Use maxDepth to control depth.

## `maxDepth` (type: `integer`):

How many levels of internal links to follow. Only used when followLinks is true.

## `maxPagesPerUrl` (type: `integer`):

Maximum pages to scrape per starting URL when following links.

## `proxyConfiguration` (type: `object`):

Proxy settings. Residential proxy recommended for protected sites.

## Actor input object example

```json
{
  "urls": [
    "https://example.com/article"
  ],
  "outputFormat": "markdown",
  "vertical": "general",
  "chunkContent": false,
  "chunkTokenSize": 512,
  "includeMetadata": true,
  "followLinks": false,
  "maxDepth": 1,
  "maxPagesPerUrl": 10
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com/article"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("conceivable_extension/llm-ready-web-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://example.com/article"] }

# Run the Actor and wait for it to finish
run = client.actor("conceivable_extension/llm-ready-web-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com/article"
  ]
}' |
apify call conceivable_extension/llm-ready-web-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=conceivable_extension/llm-ready-web-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/asd3PYJ4SnxyhNJLo/builds/bF4R6FUcWnMW8ILV5/openapi.json
