# npm Scraper - Packages, Downloads & Maintainer Leads (`scrapesage/npm-scraper`) Actor

Scrape the npm registry by keyword or package name. Get version, description, keywords, license, GitHub repo, homepage, weekly/monthly downloads, quality/popularity scores, dependencies and maintainer contact emails as B2B developer leads. No key, no login. Export JSON, CSV, Excel.

- **URL**: https://apify.com/scrapesage/npm-scraper.md
- **Developed by:** [Scrape Sage](https://apify.com/scrapesage) (community)
- **Categories:** Developer tools, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 package results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## npm Scraper — Packages, Downloads & Maintainer Leads

Search the **npm registry** by keyword or package name and get a rich, structured record for every package — **version, description, keywords, license, GitHub repo, homepage, weekly/monthly downloads, dependent counts, dependencies**, and the part other scrapers miss: **maintainer contact emails and company domains** as ready-to-use B2B developer leads.

No API key, no login, no browser — fast JSON extraction straight from the public npm registry with high reliability.

### Why this npm scraper?

The npm registry is the largest software registry in the world, and every package carries publicly published maintainer metadata. This actor merges the **search API, the registry detail API, and the downloads API** into one clean record — and turns a keyword into a scored lead list of the developers and companies behind the packages.

| Data | Typical scrapers | This actor |
|---|---|---|
| Name, version, description, keywords | ✅ | ✅ |
| GitHub repository + homepage | partial | ✅ |
| Weekly & monthly downloads | ❌ | ✅ |
| npm search relevance score | ❌ | ✅ |
| **Maintainer emails** | ❌ | ✅ |
| **Company domains** (from email addresses) | ❌ | ✅ |
| **License** | ❌ | ✅ |
| **Dependent count** (packages depending on it) | ❌ | ✅ |
| Created & last-published dates | ❌ | ✅ opt-in |
| Dependency list + version count | ❌ | ✅ opt-in |
| Deprecation status | ❌ | ✅ opt-in |
| Lead score (0–100) | ❌ | ✅ |
| Monitor mode — only new / updated packages | ❌ | ✅ |

### Use cases

- **Developer & maintainer lead generation** — sell a dev tool, SDK, API, or DevRel service? Search a niche (`stripe`, `kubernetes`, `web scraping`) and reach package maintainers at their published email, filtered to real **company domains**.
- **Developer-tool market research** — map every package in a category with downloads, dependent counts, and last-publish dates to size and track a market.
- **Supply-chain & dependency intelligence** — pull packages, their dependency lists, maintainers, and maintenance health for security and risk analysis.
- **Open-source recruiting & partnerships** — find the active maintainers behind popular packages by language, topic, or company.
- **Package monitoring** — schedule recurring runs to watch a keyword for new and freshly-published packages.

### How to use

1. [Sign up for Apify](https://console.apify.com/sign-up) — the free plan is enough to try this actor.
2. Open the **npm Scraper**, enter search queries (or package names), and click **Start**.
3. Watch packages stream into the dataset table.
4. **Export** as JSON, CSV, Excel, XML, or RSS — or pull results programmatically via the [Apify API](https://docs.apify.com/api/v2).

### Input

```json
{
    "searchQueries": ["web scraping", "keywords:cli"],
    "packageNames": ["express", "@angular/core"],
    "maxResults": 100,
    "includeDownloads": true,
    "includeDetails": true,
    "includeDependencies": false,
    "onlyWithEmail": false,
    "minMonthlyDownloads": 1000,
    "monitorMode": false
}
```

- **searchQueries** — keywords; supports npm qualifiers like `keywords:cli`, `author:sindresorhus`, `scope:angular`.
- **packageNames** — exact package names to fetch directly (always fully detailed).
- **maxResults** *(default 100)* — total packages to scrape across all queries.
- **includeDownloads** *(default true)* — add last-week and last-month download counts.
- **includeDetails** *(default false)* — fetch full registry metadata: created date, version count, deprecation, named maintainers, dependency count. (License, download counts and dependent counts ship with every search result — you do not need this switch for those.)
- **includeDependencies** *(default false)* — when details are on, output the dependency name list too.
- **onlyWithEmail / onlyWithRepository** — keep only contactable / source-linked packages.
- **minMonthlyDownloads** *(default 0)* — keep only packages above a popularity threshold.
- **monitorMode** *(default false)* — output only packages that are new or newly-versioned since the last run.

### Output

One record per package (`type: "package"`):

```json
{
    "type": "package",
    "name": "got-scraping",
    "version": "4.2.1",
    "description": "HTTP client made for scraping based on got.",
    "keywords": ["scraping", "http", "got", "crawlee"],
    "npmUrl": "https://www.npmjs.com/package/got-scraping",
    "homepage": "https://github.com/apify/got-scraping#readme",
    "repositoryUrl": "https://github.com/apify/got-scraping",
    "githubRepo": "apify/got-scraping",
    "bugsUrl": "https://github.com/apify/got-scraping/issues",
    "author": { "name": "Apify", "email": "support@apify.com" },
    "publisher": { "username": "apify-release", "email": "apify-release@apify.com" },
    "maintainers": [{ "username": "mtrunkat", "email": "marek@apify.com" }],
    "maintainerEmails": ["support@apify.com", "marek@apify.com"],
    "maintainerCount": 4,
    "companyDomains": ["apify.com"],
    "license": "ISC",
    "createdAt": "2021-03-18T12:00:00.000Z",
    "lastPublishedAt": "2026-05-30T09:14:00.000Z",
    "versionCount": 73,
    "isDeprecated": false,
    "dependenciesCount": 7,
    "weeklyDownloads": 412044,
    "monthlyDownloads": 1820551,
    "dependentsCount": 1284,
    "finalScore": 1802,
    "emailFound": true,
    "leadScore": 88,
    "searchQuery": "web scraping",
    "scrapedAt": "2026-06-27T12:00:00.000Z"
}
```

### Automate & schedule

Run this actor on autopilot and pull results into your own stack:

- **[Apify API](https://docs.apify.com/api/v2)** — start runs, fetch datasets, and manage schedules over REST.
- **[apify-client for JavaScript](https://docs.apify.com/api/client/js/)** and **[apify-client for Python](https://docs.apify.com/api/client/python/)** — official SDKs.
- **[Schedules](https://docs.apify.com/platform/schedules)** — run it daily/weekly with `monitorMode` to watch a keyword for new and updated packages, collecting only what changed.
- **[Webhooks](https://docs.apify.com/platform/integrations/webhooks)** — trigger downstream actions (CRM import, Slack alert, email sequence) the moment a run finishes.

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'MY_APIFY_TOKEN' });

const run = await client.actor('scrapesage/npm-scraper').call({
    searchQueries: ['react components'],
    maxResults: 200,
    includeDownloads: true,
    onlyWithEmail: true,
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(`Got ${items.length} packages & maintainer leads`);
```

### Integrate with any app

Connect the dataset to 5,000+ apps — no code required:

- **[Make](https://docs.apify.com/platform/integrations/make)** — multi-step automation scenarios.
- **[Zapier](https://docs.apify.com/platform/integrations/zapier)** — push new maintainer leads straight into your CRM.
- **[Slack](https://docs.apify.com/platform/integrations/slack)** — get notified when a monitored keyword gains a new package.
- **[Google Drive / Sheets](https://docs.apify.com/platform/integrations/drive)** — auto-export every run to a spreadsheet.
- **[Airbyte](https://docs.apify.com/platform/integrations/airbyte)** — pipe results into your data warehouse.
- **[GitHub](https://docs.apify.com/platform/integrations/github)** — trigger runs from commits or releases.

### Use with AI assistants (MCP)

The output is clean, LLM-ready JSON. Call this actor from Claude, ChatGPT, or any agent framework through the **[Apify MCP server](https://docs.apify.com/platform/integrations/mcp)** — ask your assistant to "find popular web-scraping npm packages with a maintainer email and over 100k monthly downloads" and let it run this scraper for you.

### Agent-ready: autonomous payments (x402 & Skyfire)

This actor is **agent-ready** — AI agents can discover it, run it, and **pay for it autonomously**, with no Apify account and no human in the loop. It uses [pay-per-event](https://docs.apify.com/platform/actors/publishing/monetize/pay-per-event) pricing and [limited permissions](https://docs.apify.com/platform/actors/development/permissions), so it qualifies for Apify's agentic-payment standards:

- **[x402](https://docs.apify.com/platform/integrations/x402)** — an open, HTTP-native payment protocol. Agents pay per run in USDC on the Base network directly through the [Apify MCP server](https://docs.apify.com/platform/integrations/mcp) — no account, no API key.
- **[Skyfire](https://docs.apify.com/platform/integrations/skyfire)** — agent-to-service payments for fully autonomous AI-agent workflows.

Building an AI agent, MCP tool, or autonomous data pipeline? This scraper is ready to plug in and pay as it goes.

### More scrapers from scrapesage

Build a complete **developer & open-source intelligence stack**:

- **[GitHub Scraper](https://apify.com/scrapesage/github-scraper)** — repos, developers, and contact leads.
- **[Hugging Face Scraper](https://apify.com/scrapesage/hugging-face-scraper)** — models, datasets, spaces, and maker leads.
- **[Chrome Web Store Scraper](https://apify.com/scrapesage/chrome-web-store-scraper)** — extensions and developer leads.
- **[WordPress Plugin & Theme Scraper](https://apify.com/scrapesage/wordpress-plugin-theme-scraper)** — installs, ratings, and author leads.
- **[Shopify App Store Scraper](https://apify.com/scrapesage/shopify-app-store-scraper)** — apps, reviews, and developer leads.
- **[Salesforce AppExchange Scraper](https://apify.com/scrapesage/salesforce-appexchange-scraper)** — ISV vendor and publisher leads.
- **[Product Hunt Scraper](https://apify.com/scrapesage/product-hunt-scraper)** — launches, makers, and leads.
- **[Website Contact Scraper](https://apify.com/scrapesage/website-contact-scraper)** — emails, phones, and socials from any website.

### Tips

- **Best lead quality**: turn on `onlyWithEmail` and set `minMonthlyDownloads` to focus on active, popular packages run by reachable maintainers — then filter the `companyDomains` column for B2B targets.
- **Search smart**: npm qualifiers narrow fast — `keywords:react`, `author:<username>`, `scope:<org>`, or `not:deprecated`.
- **Cheap vs deep**: search results already include maintainer emails, **license**, **weekly/monthly downloads** and **dependent counts**, so you get full-fat leads without `includeDetails`. Add details only when you need created dates, dependencies, version count, or deprecation.
- **Recurring monitoring**: pair [Schedules](https://docs.apify.com/platform/schedules) with `monitorMode` to track new and updated packages in a niche without re-paying for unchanged rows.

### FAQ

**Do I need an npm account or token?** No. The actor reads the public npm registry and downloads APIs — no key, account, or login required.

**Where do the maintainer emails come from?** They are the contact emails npm package authors publish in their own package metadata (`maintainers`, `author`, `publisher`). The actor surfaces and de-duplicates them, and extracts the non-personal **company domains** from them. Automated/no-reply addresses are filtered out.

**Can I export to Google Sheets, CSV, or Excel?** Yes — one click in the dataset view, or automatically on every run via the [Google Drive integration](https://docs.apify.com/platform/integrations/drive).

**How do I monitor a keyword for new packages?** Turn on `monitorMode` and create a [Schedule](https://docs.apify.com/platform/schedules). Each run outputs only packages that are new or have a new version since the previous run.

**A field is null — why?** Search results don't include created dates, dependencies, version count, or deprecation — turn on `includeDetails` for those. Some packages simply don't link a repository or homepage, have no keywords, or expose no maintainer email. `qualityScore` / `popularityScore` / `maintenanceScore` are **always** null: npm retired its per-package quality scoring and the API now returns a constant placeholder for all three, so this actor emits an honest `null` rather than a fabricated number — use `monthlyDownloads` and `dependentsCount` for popularity, and `lastPublishedAt` + `isDeprecated` for maintenance. Fields are `null` only when npm doesn't publish them.

**Is scraping npm legal?** This actor collects publicly available registry data only. You are responsible for using the data in compliance with applicable laws (GDPR/CCPA for personal data) and npm's terms.

### Need help?

Open an issue on the actor's **Issues** tab, or visit the [Apify help center](https://help.apify.com/). Feature requests are welcome — this actor is actively maintained.

# Actor input Schema

## `searchQueries` (type: `array`):

Keywords to search the npm registry, e.g. <code>web scraping</code>, <code>react components</code>, <code>stripe payments</code>. You can use npm search qualifiers like <code>keywords:cli</code>, <code>author:sindresorhus</code>, or <code>scope:angular</code>.

## `packageNames` (type: `array`):

Exact npm package names to fetch directly, e.g. <code>express</code>, <code>@angular/core</code>, <code>lodash</code>. Use this to enrich a known list of packages.

## `maxResults` (type: `integer`):

Maximum number of packages to scrape across all queries.

## `includeDownloads` (type: `boolean`):

Add last-week and last-month download counts for every package — the key popularity and lead-quality signal.

## `includeDetails` (type: `boolean`):

Fetch each package's full registry metadata: license, exact repository, created/last-published dates, version count, deprecation status, dependency list and named maintainers. Adds one request per package.

## `includeDependencies` (type: `boolean`):

When full details are on, also output the package's runtime dependency names (not just the count).

## `onlyWithEmail` (type: `boolean`):

Keep only packages that expose at least one maintainer/author/publisher email — the directly contactable developer leads.

## `onlyWithRepository` (type: `boolean`):

Keep only packages that link a source repository (usually GitHub).

## `minMonthlyDownloads` (type: `integer`):

Keep only packages with at least this many downloads in the last month (requires download counts). Leave at 0 for no filter.

## `monitorMode` (type: `boolean`):

Remember every package seen across runs and output only those that are <b>new</b> or have a <b>new version</b> since the last run. Pair with Apify Schedules to watch a niche for new and freshly-published packages — without re-paying for unchanged rows.

## `proxyConfiguration` (type: `object`):

Proxies to use. The npm registry is a public API, so the default Apify proxy is plenty — no residential needed.

## Actor input object example

```json
{
  "searchQueries": [
    "web scraping"
  ],
  "maxResults": 100,
  "includeDownloads": true,
  "includeDetails": false,
  "includeDependencies": false,
  "onlyWithEmail": false,
  "onlyWithRepository": false,
  "minMonthlyDownloads": 0,
  "monitorMode": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All scraped npm package and maintainer-lead records as JSON items in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "web scraping"
    ],
    "maxResults": 100,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapesage/npm-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": ["web scraping"],
    "maxResults": 100,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("scrapesage/npm-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "web scraping"
  ],
  "maxResults": 100,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call scrapesage/npm-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scrapesage/npm-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/jfdbNcUtQu6Sxx2to/builds/9COhbnjYQAqS9VPtf/openapi.json
