# GitHub Repo Docs Scraper for RAG & AI Agents (`ahampton83/github-docs-scraper`) Actor

Fetch documentation from GitHub repositories — READMEs, docs folders, wikis — and convert to clean, chunked markdown optimized for RAG pipelines. Use via Apify Console/API or connect as an MCP server for Claude, Cursor, and other AI agents.

- **URL**: https://apify.com/ahampton83/github-docs-scraper.md
- **Developed by:** [Aaron Hampton](https://apify.com/ahampton83) (community)
- **Categories:** AI, Developer tools, MCP servers
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## GitHub Repo Docs Scraper for RAG & AI Agents

Fetch documentation from GitHub repositories — READMEs, docs folders, and custom paths — and convert to clean, chunked markdown optimized for RAG pipelines and AI agents.

Perfect for building RAG systems over your own repos, analyzing open-source documentation, and feeding code docs to AI agents.

### Features

- **README extraction** — fetch and chunk any repo's README
- **Docs folder traversal** — recursively scan `docs/`, `documentation/`, `doc/`, and custom paths
- **Markdown chunking** — header-aware splitting with configurable chunk size and overlap
- **Rich metadata** — every chunk includes file path, line numbers, heading, and token count
- **Optional GitHub token** — supports authenticated requests (5000 req/hr vs 60 unauthenticated)
- **Dual-mode** — run as a normal Apify Actor OR connect as an MCP server for Claude, Cursor, and other AI agents
- **No browser required** — uses the GitHub REST API directly

### Use Cases

- **RAG pipelines** — chunk repo docs and ingest into vector databases for Q\&A over code documentation
- **AI agents** — let Claude/Cursor read repo documentation via MCP tools
- **Documentation analysis** — compare docs across repos, find gaps, track changes
- **Onboarding** — generate summaries of a repo's documentation for new contributors
- **Search** — keyword search across a repo's entire documentation

### Input (Normal Actor Mode)

| Field | Type | Description |
|-------|------|-------------|
| `repoUrl` | string | GitHub repo URL or `owner/repo` shorthand |
| `githubToken` | string | Optional PAT for higher rate limits |
| `includeReadme` | boolean | Fetch README (default true) |
| `includeDocsFolder` | boolean | Scan docs/ directory (default true) |
| `docsPaths` | array | Additional paths to scan |
| `maxFiles` | integer | Max files to fetch (default 50, max 500) |
| `chunkSize` | integer | Words per chunk (default 1000) |
| `chunkOverlap` | integer | Overlap between chunks (default 200) |

### Output

```json
{
  "repo": "facebook/react",
  "filePath": "README.md",
  "fileName": "README.md",
  "fileType": "md",
  "size": 5317,
  "url": "https://github.com/facebook/react/blob/main/README.md",
  "contentMarkdown": "# React · ...",
  "chunks": [
    {
      "id": "facebook_react_README_md_0",
      "text": "# React ...",
      "tokens": 465,
      "filePath": "README.md",
      "startLine": 1,
      "endLine": 10,
      "heading": "React"
    }
  ]
}
```

### MCP Tools

When connected as an MCP server, the following tools are available:

| Tool | Description |
|------|-------------|
| `get_readme` | Fetch and chunk a repo's README. |
| `get_docs` | Fetch and chunk files from a docs folder. |
| `search_repo_docs` | Fetch docs and filter by keyword. |

### Pricing

Pay-per-event. First results are free; subsequent events billed at:

| Event | Price |
|-------|-------|
| File fetched | $0.003 |
| Chunk generated | $0.001 |
| MCP tool call | $0.01 |
| Actor start | $0.00005 |

Volume discounts available at higher tiers.

# Actor input Schema

## `repoUrl` (type: `string`):

GitHub repo URL (https://github.com/owner/repo) or owner/repo shorthand.

## `githubToken` (type: `string`):

Provides 5000 req/hr instead of 60. Leave empty for public repos.

## `includeReadme` (type: `boolean`):

Fetch and process the repo's README file.

## `includeDocsFolder` (type: `boolean`):

Scan the repo's docs/ directory (and common alternatives).

## `docsPaths` (type: `array`):

Additional directory paths to scan for documentation files.

## `maxFiles` (type: `integer`):

Maximum number of files to fetch from the repository.

## `chunkSize` (type: `integer`):

Approximate words per chunk for RAG output.

## `chunkOverlap` (type: `integer`):

Words of overlap between consecutive chunks.

## `outputFormat` (type: `string`):

Return full file records or pre-chunked text for vector DBs.

## Actor input object example

```json
{
  "repoUrl": "https://github.com/microsoft/vscode",
  "includeReadme": true,
  "includeDocsFolder": true,
  "docsPaths": [],
  "maxFiles": 50,
  "chunkSize": 1000,
  "chunkOverlap": 200,
  "outputFormat": "markdown"
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "repoUrl": "https://github.com/microsoft/vscode",
    "docsPaths": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("ahampton83/github-docs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "repoUrl": "https://github.com/microsoft/vscode",
    "docsPaths": [],
}

# Run the Actor and wait for it to finish
run = client.actor("ahampton83/github-docs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "repoUrl": "https://github.com/microsoft/vscode",
  "docsPaths": []
}' |
apify call ahampton83/github-docs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=ahampton83/github-docs-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/kZsI5aXkXURo71Y0e/builds/PIhEfpnM0KnFELfQj/openapi.json
