# Techionik Website RAG MCP (`seeb/techionik-website-rag-mcp`) Actor

Prompt-first MCP-ready Apify tool for AI agents to crawl public websites and return clean RAG chunks with URLs, titles, markdown, and token estimates.

- **URL**: https://apify.com/seeb/techionik-website-rag-mcp.md
- **Developed by:** [Techionik](https://apify.com/seeb) (community)
- **Categories:** Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $15.00 / 1,000 rag chunks

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Techionik Website RAG MCP

Techionik Website RAG MCP helps AI agents build compact retrieval datasets from public websites. Users can enter a natural-language prompt, and the actor extracts URLs and crawl settings before returning clean chunks for RAG pipelines.

### What it does

- Accepts a natural-language MCP prompt for agent-friendly use.
- Also accepts structured JSON for batch/API workflows.
- Returns stable dataset records with status, warnings, source URLs, and useful business fields.
- Keeps the workflow lightweight by using public pages and clear failure records.

### Input

Use the `prompt` field when calling this as an MCP-style tool:

```json
{
    "prompt": "Create a RAG dataset from https://example.com/docs. Crawl up to 8 same-domain pages, chunk around 1600 characters, include markdown, titles, source URLs, and token estimates."
}
```

Structured fields remain available for API users who want explicit batch control.

### Local testing

Run the unit tests and project validator from this actor folder:

```bash
npm test
npm run validate
```

### Apify deployment

Before deployment, confirm the active Apify account is Techionik:

```bash
apify info
apify push
```

Marketplace publication is intentionally skipped until final business approval.

### Output

Results are written to the default Apify dataset. Each record includes a `status` field, extracted signals, warnings when something could not be fetched, and the original source input for traceability.

### Pricing

Recommended pricing: USD 1.50 per 1,000 chunks or USD 0.60 per 1,000 pages during Store billing setup. The output is immediately useful for RAG ingestion, documentation search, and AI support workflows.

### Limitations

- Works on public pages only.
- Does not bypass logins, captchas, paywalls, or site restrictions.
- Some JavaScript-heavy sites may return partial data with a warning.
- Final Apify Store billing still needs to be selected during marketplace publication.

# Actor input Schema

## `prompt` (type: `string`):

Describe the job in natural language. The actor extracts URLs, entities, and settings from this prompt, then returns structured dataset records.

## `startUrls` (type: `array`):

Optional structured fallback for API users. MCP agents should prefer prompt.

## `maxPages` (type: `integer`):

Maximum pages to fetch across all start URLs.

## `sameDomainOnly` (type: `boolean`):

Only follow links on the same hostname as each start URL.

## `chunkSize` (type: `integer`):

Approximate maximum characters per chunk.

## `chunkOverlap` (type: `integer`):

Characters of overlap between neighboring chunks.

## `includeMarkdown` (type: `boolean`):

Include a lightweight Markdown version of each chunk.

## `requestDelayMs` (type: `integer`):

Polite delay in milliseconds between public page requests.

## Actor input object example

```json
{
  "prompt": "Create a RAG dataset from https://example.com/docs. Crawl up to 8 same-domain pages, chunk around 1600 characters, include markdown, titles, source URLs, and token estimates.",
  "startUrls": [
    "https://example.com"
  ],
  "maxPages": 25,
  "sameDomainOnly": true,
  "chunkSize": 1800,
  "chunkOverlap": 200,
  "includeMarkdown": true,
  "requestDelayMs": 500
}
```

# Actor output Schema

## `results` (type: `string`):

Default dataset items produced by this MCP-compatible tool.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "prompt": "Create a RAG dataset from https://example.com/docs. Crawl up to 8 same-domain pages, chunk around 1600 characters, include markdown, titles, source URLs, and token estimates.",
    "startUrls": [
        "https://example.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("seeb/techionik-website-rag-mcp").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "prompt": "Create a RAG dataset from https://example.com/docs. Crawl up to 8 same-domain pages, chunk around 1600 characters, include markdown, titles, source URLs, and token estimates.",
    "startUrls": ["https://example.com"],
}

# Run the Actor and wait for it to finish
run = client.actor("seeb/techionik-website-rag-mcp").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "prompt": "Create a RAG dataset from https://example.com/docs. Crawl up to 8 same-domain pages, chunk around 1600 characters, include markdown, titles, source URLs, and token estimates.",
  "startUrls": [
    "https://example.com"
  ]
}' |
apify call seeb/techionik-website-rag-mcp --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=seeb/techionik-website-rag-mcp",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/VyiKZWYXlIRWU71f0/builds/sjhSwCPFpcI06gra5/openapi.json
