# 🎥 YouTube Transcript Scraper (`ethereal_wool/youtube-transcript-scraper`) Actor

Extract YouTube transcript data — name, and more. Scrape by keyword, URL or ID. Export to JSON, CSV & Excel, use the API, schedule runs and integrate. No code required.

- **URL**: https://apify.com/ethereal\_wool/youtube-transcript-scraper.md
- **Developed by:** [Jackie Chen](https://apify.com/ethereal_wool) (community)
- **Categories:** Social media, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$10.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## YouTube Transcript & Subtitle Scraper

![youtube-transcript-scraper](https://api.apify.com/v2/key-value-stores/48FFrfZnw5XqzUC1C/records/youtube-transcript-scraper__hero)

Scrape **YouTube transcripts and subtitles** for any public video. Give one or more
video IDs or URLs and get every available **caption track** — language, whether it is
auto-generated (ASR) or human-authored, whether it is translatable, and the
downloadable **timedtext URL** — together with the video's title, channel, length and
view count.

> **Unofficial.** This Actor is not affiliated with, authorized, or endorsed by YouTube
> or Google LLC. It is an independent tool that retrieves publicly available data via a
> third-party API. Use it in compliance with YouTube's Terms of Service and all
> applicable laws; you are responsible for how you use the retrieved data.

### What it does

- **Caption discovery** — for each video, lists all caption tracks YouTube exposes
  (e.g. English, Spanish, auto-generated English), with the language code, a
  human-readable name, the `kind` (`asr` = auto-generated), `isTranslatable`, and the
  `transcriptUrl` (a YouTube `timedtext` URL).
- **Video metadata** — every item also carries the parent video's `videoTitle`,
  `channel`, `channelId`, `lengthSeconds`, `viewCount` and `shortDescription`.
- **Filtering** — keep only certain languages, or only auto-generated tracks.
- **Transcript text (best effort)** — optionally tries to download and flatten the
  caption file into plain text. See the note below.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `videoIds` | string\[] | `["dQw4w9WgXcQ"]` | Video IDs or full watch / `youtu.be` / `shorts` URLs. |
| `languageCodes` | string\[] | `[]` | Keep only tracks whose language code matches (e.g. `en`, `es`). Empty = all. |
| `autoGeneratedOnly` | boolean | `false` | Keep only ASR (auto-generated) tracks. |
| `fetchTranscriptText` | boolean | `false` | Attempt to download the transcript text (best effort, see note). |
| `maxItems` | integer | `50` | Max total caption tracks across all videos. |

#### Example input

```json
{
  "videoIds": ["dQw4w9WgXcQ", "https://www.youtube.com/watch?v=jNQXAC9IVRw"],
  "languageCodes": ["en"],
  "autoGeneratedOnly": false,
  "fetchTranscriptText": true,
  "maxItems": 100
}
```

### Output

One dataset item per caption track:

```json
{
  "videoId": "dQw4w9WgXcQ",
  "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "videoTitle": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
  "channel": "Rick Astley",
  "channelId": "UCuAXFkgsw1L7xaCfnd5JJOw",
  "lengthSeconds": 213,
  "viewCount": 1779355962,
  "languageCode": "en",
  "language": "English",
  "kind": "asr",
  "isAutoGenerated": true,
  "isTranslatable": true,
  "vssId": ".en",
  "transcriptUrl": "https://www.youtube.com/api/timedtext?v=dQw4w9WgXcQ&...",
  "source": "video:dQw4w9WgXcQ"
}
```

### Notes

- **Transcript text is best-effort.** YouTube signs each `timedtext` URL against the
  IP that requested it, so a server-side download frequently returns an error. When
  `fetchTranscriptText` is enabled the Actor still tries, but `transcriptText` may come
  back empty. The `transcriptUrl` is always provided so you can fetch the caption file
  yourself (append `&fmt=json3`, `&fmt=srv3`, or `&fmt=vtt`) from the appropriate IP.
- Data is sourced live; YouTube / the upstream edge occasionally rate-limits, so the
  Actor retries transient blocks with exponential backoff.
- Video IDs are de-duplicated within a run.

### Quick start

1. Open the Actor and press **Run** — the default input works out of the box.
2. Adjust the input fields below to your target (keywords, IDs, or URLs) and set `maxItems` to cap spend.
3. Grab results from the **Dataset** tab as JSON / CSV / Excel, or pull them via the [Apify API](https://docs.apify.com/api/v2) and MCP from your own code.

No proxies to configure, no cookies to paste, no login — the Actor handles everything server-side.

### Why developers pick this transcript scraper

Transcript actors are the picks-and-shovels of the AI boom — and most charge
$10 per 1,000 videos or quietly fail on half their runs. This Actor fetches
YouTube transcripts/captions via a direct HTTP API at **$3 per 1,000 videos**,
returned as timestamped segments plus a ready-to-use plain-text field. It's
built for piping into LLMs: no HTML to clean, no SRT parsing, no browser.

### What people build with it

- **RAG knowledge bases** — index transcripts of conference talks, tutorials
  and reviews so your assistant can cite video content like documents.
- **Content repurposing** — turn long-form videos into newsletters, blog
  posts and social threads with one LLM step on top of the transcript.
- **Competitor channel analysis** — what topics, hooks and phrases do the
  top channels in your niche actually use? Transcripts answer at scale.
- **Compliance & moderation** — audit what's being said in sponsored or
  branded videos without watching hours of footage.
- **Subtitle workflows** — timestamped segments drop straight into
  translation and dubbing pipelines.
- **Research corpora** — build searchable text datasets from playlists or
  whole channels.

### Tips for better results

- Works with standard video URLs, Shorts URLs, or bare video IDs.
- Combine with [YouTube Search](https://apify.com/ethereal_wool/youtube-search-scraper)
  or [YouTube Channel Videos](https://apify.com/ethereal_wool/youtube-channel-videos-scraper)
  to discover videos first, then transcript them in bulk — a two-actor
  pipeline that turns any topic into a text corpus.
- Each segment carries `start` and `duration`, so you can deep-link to the
  exact second a phrase is spoken (`youtu.be/ID?t=123`).

***

### Why this Actor

- **Direct API, no headless browser** — fast, stable runs with nothing to babysit.
- **No login, no cookies** — we never touch your accounts, so there's no ban risk.
- **Fresh, real-time data** — every run reads the source live, not a stale cache.
- **Pay per result** — you're billed only for the rows actually delivered.
- **Structured JSON** — export to CSV, Excel, or JSON, or pull straight from the API / MCP.

### Use cases

- Build clean text corpora for LLM fine-tuning and RAG.
- Repurpose long video into blogs, summaries, and clips.
- Make video searchable and translatable at scale.
- Feed transcripts into topic modeling and keyword research.

### FAQ

**Do I need an account, cookies, or to log in anywhere?**
No. The Actor talks to a fast, direct HTTP API server-side — you just provide inputs and run it.

**How am I billed?**
Pay-per-result: a fixed price per row returned, with no separate platform/compute charge. Caps like `maxItems` keep spend predictable.

**Can I run it on a schedule or call it from my app?**
Yes — use Apify Schedules, the REST API, the JavaScript / Python clients, or the MCP server. See the **API** tab.

**Is this affiliated with YouTube?**
No. It's an independent tool that collects publicly available data. Use it in line with the platform's terms and applicable law.

### More YouTube scrapers by us

- [**YouTube Search**](https://apify.com/ethereal_wool/youtube-search-scraper) — Keyword video search · stats · channels
- [**YouTube Channel Videos**](https://apify.com/ethereal_wool/youtube-channel-videos-scraper) — All videos for a channel · stats
- [**YouTube Channel Info**](https://apify.com/ethereal_wool/youtube-channel-info-scraper) — Channel profile · subs · about
- [**YouTube Comments**](https://apify.com/ethereal_wool/youtube-comments-scraper) — Video comments + replies

Browse the full fleet → **https://apify.com/ethereal\_wool**

# Actor input Schema

## `videoIds` (type: `array`):

YouTube video IDs (e.g. dQw4w9WgXcQ) or full watch / youtu.be / shorts URLs. Each video's available caption tracks are scraped.

## `languageCodes` (type: `array`):

Only keep caption tracks whose language code matches one of these (e.g. en, es, fr). Leave empty to keep every track for each video.

## `autoGeneratedOnly` (type: `boolean`):

Keep only ASR (auto-generated) caption tracks and drop human-authored ones.

## `fetchTranscriptText` (type: `boolean`):

Also attempt to download and parse the timedtext caption file into plain transcript text. YouTube signed caption URLs are IP-locked and frequently reject server-side fetches, so this may be empty — the track URL is always returned regardless.

## `maxItems` (type: `integer`):

Maximum total number of caption tracks to push across all videos.

## `proxyConfiguration` (type: `object`):

Optional. Route the upstream API calls through an Apify Proxy to vary the source IP. Usually not needed.

## Actor input object example

```json
{
  "videoIds": [
    "dQw4w9WgXcQ"
  ],
  "languageCodes": [],
  "autoGeneratedOnly": false,
  "fetchTranscriptText": false,
  "maxItems": 50,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoIds": [
        "dQw4w9WgXcQ"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ethereal_wool/youtube-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videoIds": ["dQw4w9WgXcQ"] }

# Run the Actor and wait for it to finish
run = client.actor("ethereal_wool/youtube-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoIds": [
    "dQw4w9WgXcQ"
  ]
}' |
apify call ethereal_wool/youtube-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=ethereal_wool/youtube-transcript-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/OjXf8coxmRQngz2j5/builds/oKCBxOzSmIh7Fqgcy/openapi.json
