# GitHub Repo Metadata Enricher (`phoenix2810/github-repo-metadata-enricher`) Actor

Enrich public GitHub repository URLs with stars, forks, topics, license, activity, owner metadata, release data, and lead-scoring signals. Built for B2B lead generation, devtools research, and competitive intelligence.

- **URL**: https://apify.com/phoenix2810/github-repo-metadata-enricher.md
- **Developed by:** [Sanskar Jaiswal](https://apify.com/phoenix2810) (community)
- **Categories:** Developer tools, Lead generation, Open source
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.005 / actor start

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## GitHub Repo Metadata Enricher

Enrich public GitHub repository URLs with structured metadata, metrics, and lead-scoring signals in a single API call. Returns stars, forks, topics, license, owner info, activity timestamps, and a 0-100 lead score. Uses GitHub's public API with no authentication required.

### Use cases

- **Lead generation** - score active open-source projects by commercial signals (marketing site, license, recent activity)
- **Devtools research** - track competitor repos, adoption metrics, and integration targets
- **Market research** - map ecosystem activity and identify high-momentum projects
- **Recruiting** - find active maintainers and contributors at target companies
- **Growth teams** - build lists of repos with declared licenses and marketing sites

### Input

| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| `repoUrls` | array\<string> | yes | - | Public `github.com/<owner>/<repo>` URLs (max 50 per run) |
| `includeReadme` | boolean | no | `false` | Fetch a short README preview for each repo |
| `timeoutSeconds` | integer | no | `10` | Per-request timeout (3-30 seconds) |

#### Example input

```json
{
  "repoUrls": [
    "https://github.com/apify/crawlee",
    "https://github.com/vercel/next.js"
  ],
  "includeReadme": false,
  "timeoutSeconds": 10
}
```

### Output

Each repository produces one dataset item:

| Field | Type | Description |
|---|---|---|
| `inputUrl` | string | The GitHub URL provided |
| `fullName` | string | `owner/repo` identifier |
| `ok` | boolean | Whether enrichment succeeded |
| `checkedAt` | string | ISO timestamp |
| `owner` | object | `login`, `type` (User/Organization), `htmlUrl` |
| `repository` | object | `name`, `description`, `homepage`, `language`, `topics`, `license`, `createdAt`, `updatedAt`, `pushedAt`, `archived`, `disabled` |
| `metrics` | object | `stars`, `watchers`, `forks`, `openIssues`, `subscribers`, `sizeKb` |
| `leadSignals` | object | `score` (0-100), `signals` (array), `daysSinceLastPush` |
| `api` | object | GitHub API rate-limit status |
| `readme` | object | Optional README preview (when `includeReadme` is enabled) |
| `error` | string | Error message on failure |

#### Lead signals

The `leadSignals.score` field is a 0-100 value derived from:

| Signal | Condition | Points |
|---|---|---|
| `high_visibility` | 1,000+ stars | +35 |
| `emerging_project` | 100-999 stars | +20 |
| `developer_adoption` | 50+ forks | +15 |
| `recently_active` | Pushed within 30 days | +20 |
| `active_issue_queue` | 25+ open issues | +10 |
| `has_marketing_site` | Homepage URL set | +10 |
| `license_declared` | License explicitly declared | +5 |
| `archived_project` | Repo is archived | -40 |

#### Example output

```json
{
  "inputUrl": "https://github.com/apify/crawlee",
  "fullName": "apify/crawlee",
  "ok": true,
  "checkedAt": "2026-07-06T12:00:00.000Z",
  "owner": {
    "login": "apify",
    "type": "Organization",
    "htmlUrl": "https://github.com/apify"
  },
  "repository": {
    "name": "crawlee",
    "description": "Crawlee - A web scraping and browser automation library for Node.js",
    "homepage": "https://crawlee.dev",
    "language": "TypeScript",
    "topics": ["apify", "crawlee", "scraper", "automation"],
    "license": { "key": "apache-2.0", "name": "Apache License 2.0", "spdxId": "Apache-2.0" }
  },
  "metrics": {
    "stars": 24473,
    "forks": 1522,
    "openIssues": 176,
    "subscribers": 130
  },
  "leadSignals": {
    "score": 95,
    "signals": ["high_visibility", "developer_adoption", "recently_active", "has_marketing_site", "license_declared"],
    "daysSinceLastPush": 1
  }
}
```

### Security

- Only `github.com/<owner>/<repo>` URLs are accepted
- URLs with embedded credentials are rejected
- Uses unauthenticated public GitHub API (60 requests/hour per IP)
- No tokens, no private repos, no browser automation, no proxies
- Does not read local files or environment secrets

### Pricing

Pay per event:

| Event | Price |
|---|---|
| Actor start | $0.005 |
| Repository enriched | $0.01 |

A run enriching 10 repositories costs approximately $0.105.

### FAQ

**Do I need a GitHub token?**
No. The actor uses GitHub's unauthenticated public API. For higher volume, run on Apify infrastructure where IP rotation is automatic.

**Can I enrich private repositories?**
No. This actor only works with public repositories.

**What happens if a repo doesn't exist?**
The actor returns an item with `ok: false` and the error status. It does not crash the run.

**Can I get README content?**
Set `includeReadme: true` to fetch a 500-character README preview. This uses additional API quota.

# Actor input Schema

## `repoUrls` (type: `array`):

List of public GitHub repository URLs to enrich. Only github.com/<owner>/<repo> URLs are accepted. Up to 50 repositories per run.

## `includeReadme` (type: `boolean`):

Fetch a short README preview for each repository through GitHub's public API. Uses additional API quota but provides content context.

## `timeoutSeconds` (type: `integer`):

Timeout per GitHub API request. Lower values run faster but may fail on slow connections.

## Actor input object example

```json
{
  "repoUrls": [
    "https://github.com/apify/crawlee",
    "https://github.com/vercel/next.js"
  ],
  "includeReadme": false,
  "timeoutSeconds": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "repoUrls": [
        "https://github.com/apify/crawlee",
        "https://github.com/vercel/next.js"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("phoenix2810/github-repo-metadata-enricher").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "repoUrls": [
        "https://github.com/apify/crawlee",
        "https://github.com/vercel/next.js",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("phoenix2810/github-repo-metadata-enricher").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "repoUrls": [
    "https://github.com/apify/crawlee",
    "https://github.com/vercel/next.js"
  ]
}' |
apify call phoenix2810/github-repo-metadata-enricher --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=phoenix2810/github-repo-metadata-enricher",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XxmnRbvqS4kO8jxze/builds/WZAXkCy9IV3sHgdte/openapi.json
