# NPR Transcript Scraper — Fresh Air, Morning Edition & ATC (`scrapersdelight/npr-transcript-scraper`) Actor

Scrape full NPR transcripts — Fresh Air, Morning Edition, All Things Considered & Weekend Edition. Speaker-labeled paragraphs, full text, date, author & audio URL per story, plus a new-transcript monitor with alerts. No login or API key. $2 per 1,000 transcripts.

- **URL**: https://apify.com/scrapersdelight/npr-transcript-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Automation, AI, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 📻 NPR Transcript Scraper — Fresh Air, Morning Edition & All Things Considered

**Get the full, speaker-labeled transcript of NPR's broadcast stories — no login, no AI transcription. NPR publishes a complete transcript page for nearly every segment of Fresh Air, Morning Edition, All Things Considered and Weekend Edition, and this actor reads it: clean paragraphs, full text, speakers, date, author, section and the MP3 audio URL. Scrape one story, the latest broadcast days of any program, or page the archive back to 1991.**

Because the transcript is already published, there's no speech-to-text compute — it's fast and cheap.

***

### What does it do?

For each story (by program archive crawl or direct URL) it returns:

- 📝 **Full transcript** (plain text + `paragraphs[]`) — straight from npr.org/transcripts
- 🗣️ **Speakers** — the inline labels NPR prints (`LEILA FADEL, HOST`, `JON HAMILTON, BYLINE`, guests)
- 📅 **Broadcast date, published date, author & section**
- 🎧 **Audio URL + duration** (the segment MP3)
- 🚩 **Honest flagging** — stories whose transcript isn't published yet come back `hasTranscript: false`; nothing is synthesized

No ASR, no API key, no timestamps invented (NPR publishes none — output is paragraphs, never SRT).

***

### What data does it extract?

For every story: `storyId`, `title`, `storyUrl`, `transcriptUrl`, `program`, `episodeDate`, `publishedAt`, `author`, `section`, `speakers[]`, `paragraphs[]`, `paragraph_count`, `text`, `audioUrl`, `audioDuration`, `hasTranscript`, `is_new` (monitor), `scraped_at`.

***

### Who is it for?

- ✍️ **Journalists, researchers & students** quoting and searching radio coverage.
- 🤖 **AI / RAG builders** — dense, professionally produced news + interview transcripts, ideal retrieval/training data.
- 📰 **Newsletter writers & media monitors** tracking what NPR said about a topic, daily.
- 🎙️ **Podcast/radio analysts** studying program rundowns and guests.

***

### How to use it (step by step)

1. Click **Try for free**.
2. Pick **programs** (`fresh-air`, `morning-edition`, `all-things-considered`, `weekend-edition-saturday`, `weekend-edition-sunday`) — or paste direct **story/transcript URLs**.
3. Set **Max broadcast days** and **Max transcripts** to size the run (each archive page = 5 broadcast days; archive reaches back to 1991).
4. Click **Start**, open the **Dataset** tab to view/export.
5. *(Optional)* enable **monitorMode** + a **Schedule** to get only NEW transcripts each run, with Slack/webhook/email alerts.

#### Quick start

```json
{ "programs": ["fresh-air"], "maxEpisodes": 5, "maxStories": 10 }
```

#### Direct story lookup

```json
{ "storyUrls": ["https://www.npr.org/transcripts/nx-s1-5849937", "963319470"] }
```

***

### Input

| Field | What it does |
|-------|--------------|
| `programs` | NPR program slugs to crawl (`fresh-air`, `morning-edition`, …) |
| `storyUrls` | direct transcript/story URLs or bare story ids |
| `maxEpisodes` | recent broadcast days per program (0 = no day cap) |
| `maxStories` | hard cap on transcripts fetched per run (0 = unlimited) |
| `oldestDate` | optional `YYYY-MM-DD` floor for archive pagination |
| `includeMissingTranscripts` | also output stories with no transcript yet, flagged |
| `monitorMode`, `alertOnNewTranscript` | recurring new-transcript watcher + alerts |
| `webhookUrl`, `slackWebhookUrl`, `emailRecipients` | alert channels |
| `proxyConfiguration`, `requestConcurrency` | proxy + parallelism |

***

### Output example

```json
{
  "storyId": "nx-s1-5849937",
  "title": "Socioeconomic factors are becoming 'biologically embedded' in children's brains",
  "storyUrl": "https://www.npr.org/2026/06/11/nx-s1-5849937/child-brain-development-stress-sleep-neighborhood-economics",
  "transcriptUrl": "https://www.npr.org/transcripts/nx-s1-5849937",
  "program": "morning-edition",
  "episodeDate": "2026-06-11",
  "publishedAt": "2026-06-11",
  "author": "Jon Hamilton",
  "section": "Science",
  "speakers": ["LEILA FADEL", "JON HAMILTON"],
  "paragraphs": ["LEILA FADEL, HOST:", "New research suggests that the neighborhood a child lives in leaves a lasting imprint on their brain…", "…"],
  "paragraph_count": 18,
  "text": "LEILA FADEL, HOST:\n\nNew research suggests…",
  "audioUrl": "https://ondemand.npr.org/anon.npr-mp3/npr/me/2026/06/…mp3",
  "audioDuration": 228,
  "hasTranscript": true
}
```

Export to **JSON, CSV, Excel, HTML, or RSS**, or fetch via the **Apify API**.

***

### Monitor mode — new-transcript alerts per program

Run on a **Schedule** (e.g. every 6 hours) with `monitorMode: true`: the actor remembers every transcript it has seen (in a named, persistent store) and outputs/alerts **only the new ones**. NPR posts a story's transcript a few hours after broadcast — unseen stories are simply picked up on a later run, never faked.

```json
{ "programs": ["morning-edition", "all-things-considered"], "maxEpisodes": 3, "monitorMode": true, "slackWebhookUrl": "https://hooks.slack.com/…" }
```

***

### How much does it cost?

Pay-per-event — and with **no transcription compute**, it's cheap:

| Event | What it covers | Price |
|-------|----------------|-------|
| `lot-scraped` | each story returned | $0.004 / story |
| `lot-detail-enriched` | each transcript page fetched | $0.004 / story |
| `monitor-run-completed` | each scheduled watch run | $0.05 / run |
| `new-lot-detected` | each new transcript found | $0.02 / transcript |
| `alert-delivered` | each Slack/email/webhook push | $0.005 / alert |

That's about **$8 per 1,000 full transcripts**.

***

### Is it legal to scrape these transcripts?

This actor reads **publicly published** transcript pages on npr.org (NPR even runs a text-only site, text.npr.org). The content is NPR's (copyrighted). Scraping public pages is generally legal, but **you are responsible for your use — review NPR's terms of use and permissions policy; don't republish transcripts you're not licensed to.**

***

### FAQ

**Is there a Whisper/ASR step?**
No — NPR publishes the transcript; this actor reads it. Fast and cheap.

**Which programs work?**
Any show with an npr.org program archive: `fresh-air`, `morning-edition`, `all-things-considered`, `weekend-edition-saturday`, `weekend-edition-sunday` are verified; any `npr.org/programs/{slug}` is accepted.

**Do I get timestamps?**
No — NPR's published transcripts contain none, so the actor outputs paragraphs + full text (never fabricated SRT). You DO get the segment MP3 URL and its duration.

**Do I get speaker labels?**
Yes — NPR prints inline labels (`TERRY GROSS, HOST:`); they're kept in the paragraphs and collected into `speakers[]`.

**A story came back `hasTranscript: false` — why?**
Same-day stories gain their transcript a few hours after broadcast, and a few segments (music interludes etc.) never get one. The actor flags these honestly instead of guessing. In monitor mode they're re-checked next run.

**How far back can I go?**
The program archives paginate to 1991. Use `maxEpisodes: 0` + `oldestDate` (and a generous `maxStories`) for deep backfills.

**Both story-id formats?**
Yes — new `nx-s1-…` ids and legacy numeric ids (e.g. `963319470`) both resolve.

**How do I monitor several programs?**
List them all in `programs` — state is tracked per story id, so there's no cross-program double counting.

**How do I export?**
JSON, CSV, Excel, HTML, or RSS from the Dataset tab, or via the Apify API.

**Does it need a proxy or login?**
No login, no API key. Apify's datacenter proxy (default) is plenty — no anti-bot was observed.

***

### Feedback

Want episode-rundown mode (every segment of a broadcast day), topic filtering, or another NPR show verified? Open an issue on the actor.

# Actor input Schema

## `programs` (type: `array`):

Program slugs whose recent transcripts to scrape: fresh-air, morning-edition, all-things-considered, weekend-edition-saturday, weekend-edition-sunday (any npr.org/programs/{slug} works). Archive URLs are also accepted.

## `storyUrls` (type: `array`):

Direct NPR transcript URLs (https://www.npr.org/transcripts/{storyId}), story URLs, or bare story ids (nx-s1-5849937 or legacy 963319470). Scraped in addition to (or instead of) program archives.

## `maxEpisodes` (type: `integer`):

How many recent broadcast days (episodes) per program to crawl. Each archive page covers 5 days; the archive reaches back to 1991. 0 = no day limit (bound the run with 'Max transcripts' instead).

## `maxStories` (type: `integer`):

Hard cap on transcript pages fetched per run (0 = unlimited).

## `oldestDate` (type: `string`):

Optional YYYY-MM-DD floor — stop paginating once the archive reaches this date.

## `includeMissingTranscripts` (type: `boolean`):

Also output listed stories that have no transcript (yet), flagged hasTranscript:false. Same-day stories often gain their transcript a few hours after broadcast.

## `monitorMode` (type: `boolean`):

Recurring watcher: output/alert ONLY transcripts not seen in previous runs. Pair with a Schedule.

## `alertOnNewTranscript` (type: `boolean`):

In monitor mode, alert for each new transcript.

## `webhookUrl` (type: `string`):

POST endpoint for new-transcript alerts.

## `slackWebhookUrl` (type: `string`):

Slack incoming-webhook URL.

## `emailRecipients` (type: `array`):

Emails for the digest (via apify/send-mail).

## `proxyConfiguration` (type: `object`):

Proxy settings. Datacenter rotation is plenty — no anti-bot observed.

## `requestConcurrency` (type: `integer`):

Max parallel transcript fetches (keep modest — Akamai CDN).

## `diagnose` (type: `boolean`):

Dev only. Logs one parsed transcript, then exits.

## Actor input object example

```json
{
  "programs": [
    "fresh-air"
  ],
  "maxEpisodes": 5,
  "maxStories": 10,
  "includeMissingTranscripts": false,
  "monitorMode": false,
  "alertOnNewTranscript": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "requestConcurrency": 4,
  "diagnose": false
}
```

# Actor output Schema

## `transcripts` (type: `string`):

The dataset of NPR story transcripts (one item per story).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "programs": [
        "fresh-air"
    ],
    "maxEpisodes": 5,
    "maxStories": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/npr-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "programs": ["fresh-air"],
    "maxEpisodes": 5,
    "maxStories": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/npr-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "programs": [
    "fresh-air"
  ],
  "maxEpisodes": 5,
  "maxStories": 10
}' |
apify call scrapersdelight/npr-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scrapersdelight/npr-transcript-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/axfpXadaqg8KjZ2o9/builds/9yzmZODZSLknSRrjc/openapi.json
