# GitHub Search Scraper — Repos, Users & Issues (`ponderable_hydrometer/github-search-scraper`) Actor

Search GitHub repos, users & issues, or fetch repo details — stars, forks, topics, language, license, dates. Free GitHub API, add a token for higher rate. For lead-gen & research.

- **URL**: https://apify.com/ponderable\_hydrometer/github-search-scraper.md
- **Developed by:** [Ponderable Hydrometer](https://apify.com/ponderable_hydrometer) (community)
- **Categories:** Developer tools, Lead generation, Automation
- **Stats:** 3 total users, 2 monthly users, 95.5% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## GitHub Search Scraper — Repos, Users & Issues

**Search all of GitHub and get clean, structured rows — repositories, users, issues/PRs or code — plus full repo detail lookups by `owner/name`.** Built on GitHub's official REST API with rate-limit backoff. Add a free personal access token to run reliably at 5,000 req/h.

Perfect for developer lead-gen, open-source market research, star/trend tracking, dependency analysis, or feeding a dataset.

### What you get

Output depends on `searchType` (or the `repos` lookup):

- **Repositories** — `fullName`, `owner`, `description`, `url`, `homepage`, `language`, `stars`, `forks`, `watchers`, `openIssues`, `topics`, `license` (SPDX), `isFork`, `archived`, `defaultBranch`, `createdAt`, `updatedAt`, `pushedAt`
- **Users** — `login`, `url`, `userType`, `name`, `company`, `blog`, `location`, `bio`, `publicRepos`, `followers`, `following`
- **Issues / PRs** — `number`, `title`, `url`, `state`, `user`, `labels`, `comments`, `createdAt`, `updatedAt`, `closedAt`, `body` (a `type` of `issue` or `pull_request`)
- **Code** — `name`, `path`, `repository`, `url`, `sha`

Every row carries a `type` field so mixed exports stay unambiguous.

### Modes

- **Search** — pass a `query` in GitHub search syntax and choose `searchType` (`repositories`, `users`, `issues`, `code`). Sort and order supported. GitHub caps search itself at 1,000 results.
- **Repo detail** — pass `repos` (e.g. `["facebook/react"]`) to fetch full repository objects directly, no search needed.

### Input

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `searchType` | string | `repositories` | `repositories`, `users`, `issues`, `code` |
| `query` | string | — | GitHub search syntax, e.g. `stars:>10000 language:rust` |
| `sort` | string | best match | Repos: `stars`/`forks`/`updated`. Users: `followers`/`repositories`/`joined`. Issues: `created`/`updated`/`comments` |
| `order` | string | `desc` | `desc` or `asc` |
| `repos` | array | — | Fetch specific repos by `owner/name` |
| `token` | string (secret) | — | GitHub PAT → 5,000 req/h (read-only `public_repo` scope is enough) |
| `maxResults` | integer | `100` | Cap on results (GitHub search maxes at 1,000) |

Provide at least one of `query` or `repos`.

### Example input

```json
{
  "searchType": "repositories",
  "query": "stars:>50000 language:typescript",
  "sort": "stars",
  "maxResults": 100
}
```

Devs in a city (lead-gen):

```json
{
  "searchType": "users",
  "query": "location:berlin followers:>500",
  "sort": "followers",
  "maxResults": 50
}
```

### Example output (one repository)

```json
{
  "type": "repository",
  "id": 10270250,
  "fullName": "facebook/react",
  "name": "react",
  "owner": "facebook",
  "description": "The library for web and native user interfaces.",
  "url": "https://github.com/facebook/react",
  "homepage": "https://react.dev",
  "language": "JavaScript",
  "stars": 228000,
  "forks": 46700,
  "watchers": 228000,
  "openIssues": 970,
  "topics": ["javascript", "react", "frontend", "ui"],
  "license": "MIT",
  "isFork": false,
  "archived": false,
  "defaultBranch": "main",
  "createdAt": "2013-05-24T16:15:54Z",
  "updatedAt": "2026-07-11T09:12:00Z",
  "pushedAt": "2026-07-11T08:03:22Z"
}
```

### Why this actor

- **Four entities + detail lookups in one actor** — repos, users, issues/PRs, code, and direct `owner/name` fetches, all normalized.
- **Full GitHub search syntax** — every qualifier works (`stars:`, `language:`, `location:`, `repo:`, `state:`, …).
- **Reliable at scale** — primary and secondary rate-limit backoff; add a token for dependable 5,000 req/h throughput.
- **Flat, fair pricing** — no credit gating.

### Notes

- **Token (recommended, BYO):** without a token GitHub limits unauthenticated use to ~10 search req/min and lower per-hour caps — fine for small runs, unreliable at scale. A read-only PAT (`public_repo` scope) lifts you to 5,000 req/h. It's stored as a secret. The actor still runs keyless; it just warns and may throttle.
- GitHub search returns at most 1,000 results per query regardless of `maxResults` — narrow the query to page deeper.
- Not affiliated with or endorsed by GitHub, Inc.

### Pricing

Pay per result — **$2.00 per 1,000 results**. No subscription or platform fees; you only pay for the results you get.

### Related actors

- **Hacker News Scraper** — stories & comments for dev-community monitoring.
- **Stack Exchange Scraper** — Q\&A across Stack Overflow and sibling sites.
- **Package Registry Scraper** — npm / PyPI / crates package metadata.

# Actor input Schema

## `searchType` (type: `string`):

What to search: repositories, users, issues (incl. pull requests) or code.

## `query` (type: `string`):

GitHub search query, e.g. "stars:>10000 language:rust", "location:berlin followers:>500", "repo:apify/crawlee state:open".

## `sort` (type: `string`):

Sort field. Repos: stars, forks, updated. Users: followers, repositories, joined. Issues: created, updated, comments. Empty = best match.

## `order` (type: `string`):

"desc" or "asc".

## `repos` (type: `array`):

Fetch full details for specific repositories, e.g. \["apify/crawlee", "facebook/react"].

## `token` (type: `string`):

A GitHub PAT lifts the rate limit to 5000 req/h (vs ~10 search/min unauthenticated). Read-only 'public\_repo' scope is enough. Kept secret.

## `maxResults` (type: `integer`):

Cap on search results (GitHub search itself caps at 1000).

## Actor input object example

```json
{
  "searchType": "repositories",
  "query": "stars:>50000 language:typescript",
  "order": "desc",
  "maxResults": 100
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "stars:>50000 language:typescript"
};

// Run the Actor and wait for it to finish
const run = await client.actor("ponderable_hydrometer/github-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "stars:>50000 language:typescript" }

# Run the Actor and wait for it to finish
run = client.actor("ponderable_hydrometer/github-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "stars:>50000 language:typescript"
}' |
apify call ponderable_hydrometer/github-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=ponderable_hydrometer/github-search-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ErfIqBdlptkWyaP52/builds/e5GrJP0wIYqIM4M7f/openapi.json
