# CMS Open Data Scraper (`parseforge/data-cms-gov-scraper`) Actor

Export healthcare datasets from the Centers for Medicare & Medicaid Services Open Data portal. Pull provider directories, hospital quality, drug spending, Medicare enrollment, and 3,000+ other CMS datasets with metadata and row-level data.

- **URL**: https://apify.com/parseforge/data-cms-gov-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Business, Other, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.75 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

![ParseForge Banner](https://github.com/ParseForge/apify-assets/blob/ad35ccc13ddd068b9d6cba33f323962e39aed5b2/banner.jpg?raw=true)

## 🏥 CMS Open Data Scraper

> 🚀 **Export the U.S. Medicare and Medicaid data catalog in seconds.** Search **5,000+ healthcare datasets** by keyword, theme, or publisher, and pull the rows behind each one. No registration, no manual CSV wrangling.

The **CMS Open Data Scraper** exports the official Centers for Medicare & Medicaid Services open data catalog. Each catalog record returns **13 fields**, including dataset identifier, title, description, publisher, contact, keyword and theme tags, modified date, access level, landing page, and download link. Switch to dataset mode and the same Actor returns the rows behind any catalog entry.

The catalog covers **every public CMS dataset**: hospital cost reports, Medicare Part B and Part D drug spending, provider directories, Medicaid enrollment, nursing-home compare, marketplace open enrollment, hospital quality indicators, and thousands more. Coverage spans every state, every Medicare Administrative Contractor region, and every measurement program CMS publishes.

| 🎯 Target Audience | 💡 Primary Use Cases |
|---|---|
| Healthcare analysts, hospital finance teams, payer pricing teams, policy researchers, health-tech founders, journalists | Provider directory enrichment, drug-spend benchmarking, hospital cost analysis, quality-score lookups, payer comparison, claims research |

### 📋 What the CMS Open Data Scraper does

Two run modes in a single Actor:

- 🗂️ **Catalog mode.** List every CMS dataset matching your search term, with metadata, tags, and the download link for each one.
- 📥 **Dataset mode.** Pull the rows behind a specific dataset slug straight into your dataset.
- 🏷️ **Multi-dimensional filtering.** Restrict by keyword, theme, publisher, access level, or last-modified date.
- 🔁 **Always current.** Every run fetches the live catalog state, so your downstream dataset reflects what CMS published today.

Each catalog record carries identifiers, descriptive metadata (title, description, contact), classification tags (keyword, theme), provenance (modified date, access level), and ready-to-use links (landing page, download URL).

> 💡 **Why it matters:** the CMS catalog is one of the richest open datasets in U.S. healthcare, but the listing surface is fragmented and the per-dataset download formats vary widely. This Actor gives you a single clean shape for both the catalog and the rows underneath it.

### 📊 Data fields

Each record includes: `accessLevel`, `contactPoint`, `description`, `downloadUrl`, `identifier`, `keyword`, `landingPage`, `modified`, `publisher`, `rows`, `scrapedAt`, `theme`, `title`. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

### 🚀 How to use

1. 📝 **Sign up.** [Create a free account with $5 credit](https://console.apify.com/sign-up?fpr=vmoqkp) (takes 2 minutes).
2. 🌐 **Open the Actor.** Go to the CMS Open Data Scraper page on the Apify Store.
3. 🎯 **Set input.** Choose `catalog` or `dataset` mode, type a search term, set `maxItems`.
4. 🚀 **Run it.** Click **Start** and let the Actor collect your data.
5. 📥 **Download.** Grab your results in the **Dataset** tab as CSV, Excel, JSON, or XML.

> ⏱️ Total time from signup to downloaded dataset: **3-5 minutes.** No coding required.

### 🔗 Recommended Actors

- [**🥦 USDA FoodData Central Scraper**](https://apify.com/parseforge/usda-fooddata-central-scraper) - Official U.S. food and nutrient catalog
- [**📈 Indexmundi Scraper**](https://apify.com/parseforge/indexmundi-scraper) - Global demographic and economic indicators
- [**🔍 FINRA BrokerCheck Scraper**](https://apify.com/parseforge/finra-brokercheck-scraper) - U.S. licensed-broker reference data
- [**🗺️ Nominatim OSM Scraper**](https://apify.com/parseforge/nominatim-osm-scraper) - Geocode addresses via OpenStreetMap
- [**✈️ OurAirports Scraper**](https://apify.com/parseforge/ourairports-scraper) - Global airport reference dataset

> 💡 **Pro Tip:** browse the complete [ParseForge collection](https://apify.com/parseforge) for more reference-data scrapers.

> **⚠️ Disclaimer:** this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the Centers for Medicare & Medicaid Services or the U.S. Department of Health and Human Services. All trademarks mentioned are the property of their respective owners. Only publicly available CMS open data is collected.

### 🆘 Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our [contact form](https://tally.so/r/BzdKgA) or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our [Discord](https://parseforge.co/discord). It's the best place to get support and suggest new actors.

# Actor input Schema

## `maxItems` (type: `integer`):

How many datasets to collect per run.

## `mode` (type: `string`):

Catalog mode lists CMS datasets matching the search term. Dataset mode pulls rows from a specific dataset slug.

## `searchQuery` (type: `string`):

Keyword to filter the CMS dataset catalog. Used when Mode = Catalog. Example: 'hospital', 'medicare', 'drug spending'.

## `keyword` (type: `string`):

Match a value in the dataset's keyword tags. Example: 'quality', 'physician', 'cost report'.

## `theme` (type: `string`):

Match a value in the dataset's theme/category. Example: 'Medicare', 'Hospitals'.

## `publisher` (type: `string`):

Match the publishing organization name. Example: 'Centers for Medicare & Medicaid Services'.

## `accessLevel` (type: `string`):

Filter by data access level.

## `modifiedSince` (type: `string`):

Keep only datasets modified on or after this ISO date. Example: '2024-01-01'.

## `datasetSlug` (type: `string`):

CMS dataset identifier when Mode = Dataset. Find it in the catalog mode output as the 'identifier' field. Example: '9wzi-peqs'.

## Actor input object example

```json
{
  "maxItems": 10,
  "mode": "catalog",
  "searchQuery": "hospital"
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10,
    "mode": "catalog",
    "searchQuery": "hospital"
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/data-cms-gov-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxItems": 10,
    "mode": "catalog",
    "searchQuery": "hospital",
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/data-cms-gov-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10,
  "mode": "catalog",
  "searchQuery": "hospital"
}' |
apify call parseforge/data-cms-gov-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=parseforge/data-cms-gov-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bbztepATcBEFMk0m2/builds/Cso1NOCunZZzd2DW6/openapi.json
