# GitHub Repo Scraper — Search, Details, Releases & Trending (`thriftykiwi/github-repo-scraper`) Actor

Scrape GitHub repository data using the public REST API and trending page. Search repos by keyword, get repo details, releases, top contributors, and trending repositories. No authentication required — uses GitHub's public API endpoints and page scraping.

- **URL**: https://apify.com/thriftykiwi/github-repo-scraper.md
- **Developed by:** [Thrifty Kiwi](https://apify.com/thriftykiwi) (community)
- **Categories:** Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$5.00 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## GitHub Repositories Scraper — Search, Details, Releases, Contributors & Trending

⚡ **No authentication required** — searches public repos via the GitHub REST API and scrapes the trending page. No API tokens, no proxies, no setup.

### Quick Start

1. Pick a **mode**: Search repositories, get repo details, releases, contributors, or trending repos.
2. Enter your input — search keywords, a repo path (`owner/repo`), or a language for trending.
3. Click **Start** and get structured JSON in seconds.

#### Example: Search for machine learning repos

```json
{
  "mode": "search",
  "query": "machine learning",
  "limit": 10
}
```

#### Example: Get repo details

```json
{
  "mode": "repo",
  "owner": "microsoft",
  "repo": "vscode"
}
```

#### Example: Get trending Python repos

```json
{
  "mode": "trending",
  "language": "python",
  "limit": 25
}
```

#### Example: Get latest releases

```json
{
  "mode": "releases",
  "owner": "facebook",
  "repo": "react",
  "limit": 5
}
```

#### Example: Get top contributors

```json
{
  "mode": "contributors",
  "owner": "tensorflow",
  "repo": "tensorflow",
  "limit": 20
}
```

### Features

- **5 scraping modes**: Search, Repo Details, Releases, Contributors, Trending
- **No auth needed** — uses GitHub's public REST API and page scraping
- **No proxy required** — direct access, fast and reliable
- **Trending page scraping** — get what's hot on GitHub right now, by language
- **Pagination** — fetches up to 200 results per run
- **Structured JSON** — clean, consistent output schema with all key fields
- **Language filtering** — filter trending repos by programming language
- **Rate limit aware** — respects GitHub's public API rate limit (60 req/hr)

### Output Fields

#### Search / Repo Detail / Trending:

| Field | Description |
|-------|-------------|
| `name` | Repository name |
| `full_name` | Owner/repository |
| `description` | Repository description |
| `stars` | Star count |
| `forks` | Fork count |
| `language` | Primary programming language |
| `topics` | Repository topics/tags |
| `url` | GitHub URL |
| `created_at` | Creation date |
| `updated_at` | Last update date |
| `owner_login` | Owner username |
| `owner_type` | User or Organization |
| `license` | SPDX license identifier |
| `open_issues` | Open issue count |

#### Releases:

| Field | Description |
|-------|-------------|
| `tag_name` | Git tag (e.g., v2.1.0) |
| `name` | Release name |
| `body` | Release notes (first 500 chars) |
| `draft` | Whether it's a draft |
| `prerelease` | Whether it's a pre-release |
| `published_at` | Release date |
| `author_login` | Release author |

#### Contributors:

| Field | Description |
|-------|-------------|
| `login` | GitHub username |
| `avatar_url` | Avatar image URL |
| `contributions` | Commit count |
| `html_url` | Profile URL |

#### Trending (additional fields):

| Field | Description |
|-------|-------------|
| `stars_today` | Stars gained today |

### Use Cases

- **Competitive analysis** — Track competitor repos, their releases, and contributor activity
- **Market research** — Discover trending technologies and frameworks
- **Lead generation** — Find repos and organizations matching your ICP
- **Developer tools** — Build dashboards, alerts, and analytics on top of GitHub data
- **Academic research** — Study open-source trends, contribution patterns, and language adoption
- **Recruitment** — Find top contributors to specific technologies or projects

### Why This Scraper?

Most GitHub scrapers require auth tokens, use heavyweight browser automation, or only do one thing. This scraper:

- **No token needed** — uses the public REST API for instant results
- **All 5 modes in one** — search, details, releases, contributors, and trending
- **Trending page scraping** — unique feature not available via the REST API
- **Clean defaults** — click "Start" with defaults and get real results immediately
- **Lightweight** — HTTP-only, no browser overhead

### Rate Limits

This Actor uses GitHub's public REST API which has a rate limit of **60 requests per hour** per IP. For higher volumes, provide a GitHub personal access token — but even without one, the default limits are sufficient for most use cases.

### Pricing

This Actor uses **Pay-Per-Event** pricing. You only pay for results returned, not for API calls or pagination overhead.

### About the Publisher

Built by [Thrifty Kiwi](https://apify.com/thriftykiwi), an Apify Creator publishing high-quality, zero-hassle data extraction tools.

# Actor input Schema

## `mode` (type: `string`):

What type of GitHub data to scrape.

## `query` (type: `string`):

Keyword or phrase to search GitHub repositories (only used in 'search' mode).

## `owner` (type: `string`):

GitHub username or organization. Required for repo, releases, and contributors modes.

## `repo` (type: `string`):

Repository name (without owner). Required for repo, releases, and contributors modes.

## `language` (type: `string`):

Filter trending repos by programming language (lowercase). Leave empty for all languages. Only used in 'trending' mode. Examples: python, javascript, rust, go.

## `spokenLanguage` (type: `string`):

Spoken language filter for trending (e.g., 'en' for English). Leave empty for default. Only used in 'trending' mode.

## `limit` (type: `integer`):

Maximum number of results to return (1-200).

## Actor input object example

```json
{
  "mode": "search",
  "query": "machine learning",
  "owner": "microsoft",
  "repo": "vscode",
  "language": "",
  "spokenLanguage": "",
  "limit": 10
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "search",
    "query": "machine learning",
    "owner": "microsoft",
    "repo": "vscode",
    "language": "",
    "spokenLanguage": "",
    "limit": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("thriftykiwi/github-repo-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "search",
    "query": "machine learning",
    "owner": "microsoft",
    "repo": "vscode",
    "language": "",
    "spokenLanguage": "",
    "limit": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("thriftykiwi/github-repo-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "search",
  "query": "machine learning",
  "owner": "microsoft",
  "repo": "vscode",
  "language": "",
  "spokenLanguage": "",
  "limit": 10
}' |
apify call thriftykiwi/github-repo-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=thriftykiwi/github-repo-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/cm2QpWnVthQ8ukCUV/builds/CVIpmKaUZ5akZaVBN/openapi.json
