# YouTube Research Scraper - Videos, Channels & Search (`lentic_clockss/youtube-research-scraper`) Actor

Collect YouTube video and channel research data for content analysis, competitor monitoring, and lead research. Export structured metadata for automation workflows.

- **URL**: https://apify.com/lentic\_clockss/youtube-research-scraper.md
- **Developed by:** [kane liu](https://apify.com/lentic_clockss) (community)
- **Categories:** Social media, Videos, Developer tools
- **Stats:** 10 total users, 2 monthly users, 92.9% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 base video rows

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## YouTube Research & Transcript Scraper

Search YouTube, export channel video lists, enrich selected videos, and collect transcripts without setting up the YouTube Data API.

This Apify Actor is built for YouTube research workflows where you do not want to scrape everything at the most expensive level. Start with broad discovery, shortlist the videos that matter, then run metadata enrichment or transcript extraction only on that smaller set.

### Best for

- YouTube keyword research and topic mapping
- competitor and creator channel monitoring
- content audits for brands, agencies, and media teams
- transcript collection for LLM, RAG, summarization, and qualitative research pipelines
- building structured YouTube datasets from search results, channel pages, and known video URLs

### How it works

The Actor accepts three input types. You can use one, two, or all three in the same run.

| Input | Use it when | Result source |
| --- | --- | --- |
| `searchQueries` | You want to discover videos by topic or keyword | YouTube search results |
| `channelUrls` | You want recent videos from one or more channels | Channel `/videos` pages with browse continuation, plus RSS fallback when needed |
| `videoUrls` | You already know the exact videos to process | Direct video metadata paths with fallback metadata extraction |
| `includeTrending` | You want popular videos for a market (`gl`) | Topic hubs (gaming, sports, news, podcasts, live, learning, fashion) — classic `/feed/trending` was removed by YouTube |

Rows are deduplicated by `videoId`, so the same video is only pushed once even if it appears in multiple inputs.

At least one of `searchQueries`, `channelUrls`, `videoUrls`, or `includeTrending` must be set. Empty input is rejected so the Actor does not create a misleading dataset row or charge for a helper item.

#### Comments (full coverage)

Set `includeComments: true` on a shortlist (`videoUrls` recommended). The worker paginates InnerTube `/next` for top-level comments and reply threads (modern `commentEntityPayload` text). Use `maxComments: 0` / `maxRepliesPerComment: 0` for exhaustive crawls within safety caps (20k tops / 2k replies per thread). `includeRelated` attaches watch-page related videos on each row.

### More Actors like this

Looking for another **social / video** scraper, or a specialized Actor outside YouTube? Use a dedicated Actor when one exists — structured fields, better coverage, usually lower cost.

#### Similar social & content Actors

- [YouTube Shorts Scraper](https://apify.com/lentic_clockss/youtube-shorts-scraper)
- [TikTok Scraper](https://apify.com/lentic_clockss/tiktok-scraper)
- [Reddit Scraper](https://apify.com/lentic_clockss/reddit-scraper)
- [Hacker News Scraper](https://apify.com/lentic_clockss/hacker-news-scraper)

#### Prefer another specialized scraper?

**Jobs & Freelance**

- [LinkedIn Jobs Scraper](https://apify.com/lentic_clockss/linkedin-jobs-scraper)
- [Indeed Jobs Scraper](https://apify.com/lentic_clockss/indeed-jobs-scraper)
- [Upwork Jobs Scraper](https://apify.com/lentic_clockss/upwork-jobs-scraper)
- [Glassdoor Scraper](https://apify.com/lentic_clockss/glassdoor-scraper)
- [Fiverr Gigs Scraper](https://apify.com/lentic_clockss/fiverr-programming-tech-gigs-scraper)
- [Bayt Jobs Scraper](https://apify.com/lentic_clockss/bayt-scraper)

**E-commerce**

- [Walmart Product Scraper](https://apify.com/lentic_clockss/walmart-scraper)
- [Amazon Search Scraper](https://apify.com/lentic_clockss/amazon-search-results-collector)
- [Shopee Search Scraper](https://apify.com/lentic_clockss/shopee-search-scraper)
- [Etsy Scraper](https://apify.com/lentic_clockss/etsy-scraper)
- [SHEIN Product Scraper](https://apify.com/lentic_clockss/shein-scraper)
- [Temu Product Scraper](https://apify.com/lentic_clockss/temu-scraper)
- [Target Product Scraper](https://apify.com/lentic_clockss/target-scraper)
- [Allegro Scraper](https://apify.com/lentic_clockss/allegro-scraper)

**Real Estate**

- [Zillow & Zumper Scraper](https://apify.com/lentic_clockss/us-real-estate-scraper)
- [Realtor.com Scraper](https://apify.com/lentic_clockss/realtor-com-scraper)
- [Apartments.com Rental Scraper](https://apify.com/lentic_clockss/apartments-com-rental-scraper)
- [Rightmove Scraper](https://apify.com/lentic_clockss/rightmove-property-scraper)
- [Idealista Scraper](https://apify.com/lentic_clockss/idealista-scraper)
- [realestate.com.au Scraper](https://apify.com/lentic_clockss/realestate-com-au-scraper)

**Travel & Stays**

- [Booking.com Hotels Scraper](https://apify.com/lentic_clockss/booking-hotels-scraper)
- [Airbnb Listings Scraper](https://apify.com/lentic_clockss/airbnb-listings-scraper)
- [Expedia Scraper](https://apify.com/lentic_clockss/expedia-scraper)
- [TripAdvisor Scraper](https://apify.com/lentic_clockss/tripadvisor-scraper)

**Ads Intelligence**

- [Facebook Ad Library Scraper](https://apify.com/lentic_clockss/facebook-ad-library-scraper)
- [TikTok Ads Scraper](https://apify.com/lentic_clockss/tiktok-ads-top-ads-actor)

**Local & Maps**

- [Google Maps Scraper](https://apify.com/lentic_clockss/google-maps-scraper)

**General Tools**

- [Stealth Web Scraper](https://apify.com/lentic_clockss/stealth-web-scraper)
- [Email Risk Validator](https://apify.com/lentic_clockss/email-risk-validator)
- [Phone Number Intelligence](https://apify.com/lentic_clockss/phone-number-intelligence)

→ See the full catalog in [Related Actors](#related-actors) below, or browse [apify.com/lentic\_clockss](https://apify.com/lentic_clockss).

***

### How to use (no code required)

1. Click **"Try for Free"** at the top of this page
2. Add at least one input: `searchQueries`, `channelUrls`, `videoUrls`, and/or turn on `includeTrending`
3. Keep `scrapeDetails` / `includeTranscript` / `includeComments` off for cheap discovery; turn them on only for a shortlist
4. Click **Start** — rows appear in the Dataset tab
5. Download as **JSON, CSV, or Excel**, or call the Standby API for small interactive requests

**Tip:** discover broadly first, then enrich or pull transcripts only for the videos you actually need — that keeps cost and runtime down.

***

### Recommended workflow

#### 1. Discover videos cheaply

Use `searchQueries` or `channelUrls` first. Keep `scrapeDetails` and `includeTranscript` off while you are still exploring.

```json
{
  "searchQueries": ["ai workflow automation", "youtube competitor analysis"],
  "maxResults": 50
}
```

This gives you a clean shortlist with titles, URLs, channels, thumbnails, rough publish text, durations, view counts when available, and descriptions when present in the search result.

#### 2. Review and shortlist

Filter the dataset outside the Actor. Pick only the videos you actually need for deeper work.

Useful shortlist signals:

- topic relevance from `title` and `description`
- creator or company from `channelName`
- popularity from `viewCount`
- freshness from `publishedText` or `publishedAt`
- video length from `duration` or `durationSeconds`

#### 3. Enrich selected videos

Use `videoUrls` with `scrapeDetails` when you need stronger metadata for specific videos.

```json
{
  "videoUrls": [
    "https://www.youtube.com/watch?v=XVv6mJpFOb0",
    "https://youtu.be/dQw4w9WgXcQ"
  ],
  "scrapeDetails": true
}
```

`scrapeDetails` may improve or fill:

- `publishedAt`
- `category`
- `description`
- `viewCount`

It is best used after shortlisting because it performs extra requests per video.

#### 4. Collect transcripts only when needed

Use `includeTranscript` for videos where you actually need text, timestamps, or LLM-ready content.

```json
{
  "videoUrls": [
    "https://www.youtube.com/watch?v=XVv6mJpFOb0"
  ],
  "scrapeDetails": true,
  "includeTranscript": true,
  "transcriptLanguage": "en"
}
```

When a transcript is available, the row includes timestamped transcript segments and a combined plain-text transcript. If YouTube does not provide captions for the video, or the captions cannot be fetched, the Actor still returns the video row without transcript fields.

### Input reference

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `searchQueries` | array of strings | empty | YouTube search keywords. Each query runs separately and can return up to `maxResults` videos. Best for discovery and SEO or market research. |
| `channelUrls` | array of strings | empty | YouTube channel inputs. Supports `@handle`, UC channel IDs, and common `youtube.com` channel, c, and user URLs. Returns recent public videos; RSS fallback is used when the channel page does not expose rows. |
| `videoUrls` | array of strings | empty | Exact YouTube videos to process. Supports 11-character video IDs and common watch, shorts, embed, live, and `youtu.be` URL formats. Best for enrichment and transcripts. |
| `maxResults` | integer | `50` | Maximum videos per search query or channel. It does not multiply direct `videoUrls`; each provided video URL is processed once. |
| `scrapeDetails` | boolean | `false` | Fetches richer metadata for each row. Use on shortlists or smaller runs. |
| `includeTranscript` | boolean | `false` | Attempts transcript extraction for each video. Use on targeted runs because this is the heaviest mode. |
| `transcriptLanguage` | string | `en` | Preferred transcript language code, such as `en`, `es`, `fr`, `de`, `ja`, or `pt`. If that language is unavailable, the Actor can fall back to the first available caption track. |
| `includeComments` | boolean | `false` | Full comment pagination (+ nested replies). Prefer with `videoUrls`. |
| `maxComments` | integer | `0` | Max top-level comments; `0` = all (cap 20000). |
| `includeCommentReplies` | boolean | `true` | Expand reply threads with full pagination. |
| `maxRepliesPerComment` | integer | `0` | Max replies per thread; `0` = all (cap 2000). |
| `commentSort` | string | `top` | `top` or `newest`. |
| `includeRelated` | boolean | `false` | Attach related/recommended videos per watch page. |
| `maxRelated` | integer | `20` | Cap related videos per source video. |
| `includeTrending` | boolean | `false` | Discover popular videos via topic hubs for `gl`. |
| `trendingMaxResults` | integer | `50` | Cap for topic-hub discovery. |
| `gl` / `hl` | string | `US` / `en` | Market / UI language for search, comments, and topic hubs. |

### Input examples

#### Search by keyword

```json
{
  "searchQueries": ["supply chain automation"],
  "maxResults": 25
}
```

#### Export latest channel videos

```json
{
  "channelUrls": ["https://www.youtube.com/@freecodecamp"],
  "maxResults": 100
}
```

#### Process known videos

```json
{
  "videoUrls": [
    "https://www.youtube.com/watch?v=XVv6mJpFOb0",
    "https://youtu.be/PXMJ6FS7llk"
  ],
  "scrapeDetails": true
}
```

#### Transcript run for a shortlist

```json
{
  "videoUrls": [
    "https://www.youtube.com/watch?v=XVv6mJpFOb0"
  ],
  "includeTranscript": true,
  "transcriptLanguage": "en"
}
```

#### Mixed discovery run

```json
{
  "searchQueries": ["ai sales outreach"],
  "channelUrls": ["https://www.youtube.com/@HubSpot"],
  "maxResults": 30
}
```

### Output fields

Each dataset item is one YouTube video row. The Actor does not write helper rows for empty input.

#### Core fields

| Field | Type | Description |
| --- | --- | --- |
| `recordVersion` | string | Output contract version, currently `1.0`. |
| `enrichmentLevel` | string | `base`, `detail`, or `transcript`. Shows how far the row was enriched. |
| `videoId` | string | YouTube video ID. |
| `title` | string | Video title. |
| `url` | string | Canonical YouTube watch URL. |
| `channelName` | string | Channel or author name when available. |
| `channelId` | string | YouTube channel ID when available. |
| `channelUrl` | string | Channel URL when available. |
| `viewCount` | integer | View count when available. May be `0` when the source does not expose it. |
| `duration` | string | Human-readable duration from listing pages when available. |
| `durationSeconds` | integer | Duration in seconds when available. |
| `publishedText` | string | Relative publish text from listing pages, such as `2 weeks ago`, when available. |
| `publishedAt` | string | Publish date when available. Detail mode can improve this field. |
| `description` | string | Search snippet, RSS description, or fuller video description depending on source and enrichment. |
| `thumbnailUrl` | string | Video thumbnail URL. |
| `category` | string | Video category when detail metadata is available. |
| `isLive` | boolean | Whether the source marks the video as live content. |
| `source` | string | Source path used for the row: `search`, `channel`, or `detail`. |
| `scrapedAt` | string | ISO timestamp when the row was created. |

#### Transcript fields

Transcript fields appear only when `includeTranscript` is true and captions are successfully returned.

| Field | Type | Description |
| --- | --- | --- |
| `transcript` | array | Timestamped caption segments. Each segment has `text`, `start`, and `duration`. |
| `transcriptLanguage` | string | Language code of the transcript actually returned. |
| `transcriptText` | string | Full transcript joined into one plain-text string. |

Example transcript segment:

```json
{
  "text": "Welcome back to the channel.",
  "start": 12.4,
  "duration": 3.2
}
```

### Standby API

The Actor includes a Standby API for small interactive requests. The same validation rules apply as normal runs.

| Endpoint | Method | Use |
| --- | --- | --- |
| `/` | `GET` | Readiness check |
| `/search?query=python%20automation&maxResults=10` | `GET` | Search videos |
| `/channel?url=https://www.youtube.com/@freecodecamp&maxResults=10` | `GET` | List recent channel videos |
| `/video?url=XVv6mJpFOb0` | `GET` | Fetch one direct video |
| `/run` | `POST` | Run the normal Actor input JSON through Standby |

### Limits and practical notes

- Transcripts are not guaranteed. They depend on whether YouTube exposes captions for the video and whether those captions can be fetched.
- `includeTranscript` can still return a valid video row without transcript fields.
- Search and channel rows may have lighter metadata than direct detail rows.
- `maxResults` applies per search query and per channel URL.
- Channel scraping works best with public channels and common YouTube URL formats.
- Very large transcript runs are slower and more expensive than discovery runs. Shortlist first when possible.
- YouTube page structure and availability can change. If a source path fails for a specific video or channel, try the most direct input type, especially `videoUrls` for known videos.

### Pricing model

The Actor uses tiered pay-per-event charging with these event keys:

1. `apify-default-dataset-item` — base video rows for search and channel discovery
2. `youtube-video-detail` — detailed video rows when detail enrichment succeeds
3. `youtube-video-transcript` — transcript-ready rows when transcript extraction succeeds

Charge tier follows what was actually delivered for each row. Discover broadly first, then run detail or transcript modes only on a shortlist.

Check the Apify Store pricing panel for the current event prices before running large jobs.

### Local tests

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt pytest
pytest -q
```

### Why use this Actor

This Actor is focused on research, not just bulk scraping. It separates discovery, detail enrichment, and transcript extraction so you can control speed, dataset size, and cost.

Use it when you need structured YouTube data for market research, creator research, competitor monitoring, content strategy, or LLM-ready transcript workflows without maintaining your own YouTube scraping stack.

### Operational hardening

This Actor emits structured progress logs so long runs are easier to diagnose from Apify logs and Insights:

- `progress_event=run_input_ready` after the input is normalized.
- `progress_event=source_start` / `source_done` / `source_error` for search, channel, and video sources.
- `progress_event=detail_enrich_start` / `detail_enrich_done` for optional video detail enrichment.
- `progress_event=transcript_start` / `transcript_done` for optional transcript extraction.
- `progress_event=row_push_start` / `row_push_done` for dataset writes and billing-event boundaries.
- `progress_event=run_summary_ready` before `RUN_SUMMARY` is written.

HTTP requests use curl\_cffi Chrome impersonation plus coherent browser headers, consent cookies, InnerTube client context, and YouTube-specific JSON headers to reduce obvious datacenter/client-fingerprint mismatches while keeping the Actor lightweight and API-first.

### Live-view web server OpenAPI schema

This Actor includes a real Actor Standby / Live-view web server schema at:

```text
.actor/openapi.json
```

The schema is published and validated through `.actor/actor.json`:

```json
{
  "usesStandbyMode": true,
  "webServerSchema": "./openapi.json"
}
```

Documented Standby endpoints:

- `GET /` - service information and Apify readiness-probe response
- `GET /health` - health check
- `GET /input-example` - quick YouTube research request examples
- `GET /openapi.json` - returns the OpenAPI document packaged with the Actor
- `GET /search` and `POST /search` - runs a bounded YouTube search
- `GET /channel` - scans a YouTube channel
- `GET /video` - processes one exact video URL or ID
- `POST /run` - runs the Actor with the full YouTube Research Scraper input contract

For low-cost validation, use `POST /search` with a small `maxResults` value and `includeTranscript: false`.

***

### Related Actors

All **77** public Actors from \[[lentic\_clockss](https://apify.com/lentic_clockss)]. Click a name to open the Store detail page.

#### Jobs & Freelance

- [LinkedIn Jobs Scraper](https://apify.com/lentic_clockss/linkedin-jobs-scraper)
- [Bayt Jobs Scraper](https://apify.com/lentic_clockss/bayt-scraper)
- [Fiverr Gigs Scraper](https://apify.com/lentic_clockss/fiverr-programming-tech-gigs-scraper)
- [Freelancer.com Scraper](https://apify.com/lentic_clockss/freelancer-scraper)
- [Glassdoor Scraper](https://apify.com/lentic_clockss/glassdoor-scraper)
- [Himalayas Jobs Scraper](https://apify.com/lentic_clockss/himalayas-jobs-scraper)
- [Indeed Jobs Scraper](https://apify.com/lentic_clockss/indeed-jobs-scraper)
- [Jobicy Remote Jobs Scraper](https://apify.com/lentic_clockss/jobicy-remote-jobs-scraper)
- [RemoteOK Jobs Scraper](https://apify.com/lentic_clockss/remoteok-all-jobs-scraper)
- [SEEK Jobs Scraper](https://apify.com/lentic_clockss/seek-scraper)
- [Upwork Jobs Scraper](https://apify.com/lentic_clockss/upwork-jobs-scraper)

#### Real Estate

- [Zillow & Zumper Scraper](https://apify.com/lentic_clockss/us-real-estate-scraper)
- [Realtor.com Scraper](https://apify.com/lentic_clockss/realtor-com-scraper)
- [99.co Scraper](https://apify.com/lentic_clockss/ninetynine-co-sg-scraper)
- [Realtor.com Agents Scraper](https://apify.com/lentic_clockss/realtor-com-agents-scraper)
- [Apartments.com Rental Scraper](https://apify.com/lentic_clockss/apartments-com-rental-scraper)
- [Bayut Scraper](https://apify.com/lentic_clockss/bayut-uae-scraper)
- [Craigslist Housing Scraper](https://apify.com/lentic_clockss/craigslist-housing-scraper)
- [Daft.ie Scraper](https://apify.com/lentic_clockss/daft-property-scraper)
- [Dot Property Scraper](https://apify.com/lentic_clockss/dot-property-th-scraper)
- [FINN.no Scraper](https://apify.com/lentic_clockss/finn-scraper)
- [Funda Scraper](https://apify.com/lentic_clockss/funda-scraper)
- [Hepsiemlak Scraper](https://apify.com/lentic_clockss/hepsiemlak-scraper)
- [Idealista Scraper](https://apify.com/lentic_clockss/idealista-scraper)
- [Immobiliare.it Scraper](https://apify.com/lentic_clockss/immobiliare-property-scraper)
- [ImmoScout24 Scraper](https://apify.com/lentic_clockss/immobilienscout24-scraper)
- [Naver Land Scraper](https://apify.com/lentic_clockss/naver-land-scraper)
- [OpenSooq Scraper](https://apify.com/lentic_clockss/opensooq-real-estate-scraper)
- [Otodom Scraper](https://apify.com/lentic_clockss/otodom-scraper)
- [Property Finder Scraper](https://apify.com/lentic_clockss/property-finder-uae-scraper)
- [PropertyGuru Scraper](https://apify.com/lentic_clockss/propertyguru-sg-scraper)
- [realestate.com.au Scraper](https://apify.com/lentic_clockss/realestate-com-au-scraper)
- [Realtor.ca Scraper](https://apify.com/lentic_clockss/realtor-ca-scraper)
- [Rightmove Scraper](https://apify.com/lentic_clockss/rightmove-property-scraper)
- [SeLoger Scraper](https://apify.com/lentic_clockss/seloger-property-scraper)
- [SUUMO Scraper](https://apify.com/lentic_clockss/suumo-property-scraper)
- [Zillow Group Scraper](https://apify.com/lentic_clockss/zillow-group-scraper)

#### E-commerce

- [Shopee Search Scraper](https://apify.com/lentic_clockss/shopee-search-scraper)
- [E-commerce Scraper](https://apify.com/lentic_clockss/ecommerce-scraper)
- [1688 Global Product Search Scraper](https://apify.com/lentic_clockss/1688-global-scraper)
- [Allegro Scraper](https://apify.com/lentic_clockss/allegro-scraper)
- [Amazon Search Scraper](https://apify.com/lentic_clockss/amazon-search-results-collector)
- [ASOS Product Scraper](https://apify.com/lentic_clockss/asos-scraper)
- [Cdiscount Product Scraper](https://apify.com/lentic_clockss/cdiscount-scraper)
- [Costco Product Scraper](https://apify.com/lentic_clockss/costco-scraper)
- [Coupang Product Scraper](https://apify.com/lentic_clockss/coupang-scraper)
- [Etsy Scraper](https://apify.com/lentic_clockss/etsy-scraper)
- [Lazada Scraper](https://apify.com/lentic_clockss/lazada-ph-search-results-collector)
- [MercadoLibre Scraper](https://apify.com/lentic_clockss/mercadolibre-scraper)
- [Mercari Japan Scraper](https://apify.com/lentic_clockss/mercari-scraper)
- [Rakuten Japan Scraper](https://apify.com/lentic_clockss/rakuten-scraper)
- [SHEIN Product Scraper](https://apify.com/lentic_clockss/shein-scraper)
- [Target Product Scraper](https://apify.com/lentic_clockss/target-scraper)
- [Temu Product Scraper](https://apify.com/lentic_clockss/temu-scraper)
- [Walmart Product Scraper](https://apify.com/lentic_clockss/walmart-scraper)

#### Travel & Stays

- [Booking.com & Airbnb Scraper](https://apify.com/lentic_clockss/booking-airbnb-scraper)
- [Agoda Scraper](https://apify.com/lentic_clockss/agoda-scraper)
- [Airbnb Listings Scraper](https://apify.com/lentic_clockss/airbnb-listings-scraper)
- [Booking.com Hotels Scraper](https://apify.com/lentic_clockss/booking-hotels-scraper)
- [Despegar Scraper](https://apify.com/lentic_clockss/despegar-scraper)
- [Expedia Scraper](https://apify.com/lentic_clockss/expedia-scraper)
- [Traveloka Scraper](https://apify.com/lentic_clockss/traveloka-scraper)
- [Travelstart Flights Scraper](https://apify.com/lentic_clockss/travelstart-scraper)
- [Trip.com Scraper](https://apify.com/lentic_clockss/trip-com-scraper)
- [TripAdvisor Scraper](https://apify.com/lentic_clockss/tripadvisor-scraper)

#### Social & Content

- [TikTok Scraper](https://apify.com/lentic_clockss/tiktok-scraper)
- [Reddit Scraper](https://apify.com/lentic_clockss/reddit-scraper)
- [YouTube Shorts Scraper](https://apify.com/lentic_clockss/youtube-shorts-scraper)
- [YouTube Research Scraper](https://apify.com/lentic_clockss/youtube-research-scraper)
- [Hacker News Scraper](https://apify.com/lentic_clockss/hacker-news-scraper)

#### Ads Intelligence

- [Facebook Ad Library Scraper](https://apify.com/lentic_clockss/facebook-ad-library-scraper)
- [Google Ads Transparency VN](https://apify.com/lentic_clockss/google-ads-transparency-center-vn)
- [TikTok Ads Scraper](https://apify.com/lentic_clockss/tiktok-ads-top-ads-actor)

#### Local & Maps

- [Google Maps Scraper](https://apify.com/lentic_clockss/google-maps-scraper)

#### General Tools

- [Stealth Web Scraper](https://apify.com/lentic_clockss/stealth-web-scraper)
- [Email Risk Validator](https://apify.com/lentic_clockss/email-risk-validator)
- [Phone Number Intelligence](https://apify.com/lentic_clockss/phone-number-intelligence)

→ Browse the full profile: [apify.com/lentic\_clockss](https://apify.com/lentic_clockss)

# Actor input Schema

## `runMode` (type: `string`):

Use real for live YouTube collection via the VPS worker. Use fixture for offline smoke tests (no live YouTube traffic).

## `workerBaseUrl` (type: `string`):

Optional HTTPS override for the scrape worker. Defaults to Actor env WORKER\_BASE\_URL or https://yt.opendata.best.

## `searchQueries` (type: `array`):

YouTube keywords to search. Each query runs independently and returns up to Max results videos. Use this for topic discovery, SEO research, competitor research, and shortlist building before heavier detail or transcript runs.

## `channelUrls` (type: `array`):

Public YouTube channels to scan for recent videos. Supports @handle, UC channel IDs, and common youtube.com channel, c, and user URLs. Use this for creator, brand, media, or competitor channel monitoring.

## `videoUrls` (type: `array`):

Exact YouTube videos to process. Supports 11-character video IDs and common watch, youtu.be, shorts, embed, and live URL formats. Best for targeted metadata enrichment and transcript extraction after you already have a shortlist.

## `maxResults` (type: `integer`):

Maximum number of videos to return for each search query or each channel URL. This does not limit direct Video URLs; every listed video URL is processed once. Keep this lower when using Scrape video details or Include transcript.

## `gl` (type: `string`):

YouTube market / region code for search ranking and InnerTube context (ISO-3166 alpha-2), e.g. US, GB, DE, JP, BR, IN.

## `hl` (type: `string`):

YouTube UI/search language, e.g. en, de, ja, pt-BR. Independent from transcriptLanguage.

## `proxyCountry` (type: `string`):

Optional residential egress country for detail/transcript (and search when Use proxy for search is on). Defaults to gl when gl/hl are set; otherwise US for enrichment-only runs.

## `uploadDate` (type: `string`):

YouTube search upload-date filter.

## `resultType` (type: `string`):

YouTube search type filter. Use short for short-form bias.

## `duration` (type: `string`):

YouTube search duration filter (short under ~4m, long over ~20m).

## `channelTab` (type: `string`):

Which channel tab to scrape: videos, shorts, streams, or all (split budget across tabs).

## `includeChannelAbout` (type: `boolean`):

When scraping channels, also extract about/stats into RUN diagnostics (subscriber text, country, description).

## `useProxyForSearch` (type: `boolean`):

Route search/channel HTML through residential proxy. Off by default (datacenter InnerTube/HTML); enable for geo-sensitive SERP.

## `concurrency` (type: `integer`):

Max parallel detail/transcript/comment enrichments (1–16). Higher is faster but uses more proxy/egress.

## `includeComments` (type: `boolean`):

Fetch comments for each video via InnerTube /next with full pagination. Nest comments (+ replies) on each video row. Use with videoUrls or small shortlists for complete coverage.

## `maxComments` (type: `integer`):

Max top-level comments per video. 0 = fetch all until exhausted (hard safety cap 20000).

## `includeCommentReplies` (type: `boolean`):

Expand reply threads for each top-level comment (full pagination per thread).

## `maxRepliesPerComment` (type: `integer`):

Max replies per top-level comment. 0 = all until exhausted (hard safety cap 2000).

## `commentSort` (type: `string`):

YouTube comment sort order.

## `includeRelated` (type: `boolean`):

Attach related/recommended videos from the watch /next response onto each video row.

## `maxRelated` (type: `integer`):

Max related videos to attach per source video.

## `includeTrending` (type: `boolean`):

Discover popular videos for the selected market (gl). YouTube removed classic /feed/trending; this uses topic hubs (gaming, sports, news, podcasts, live, learning, fashion) and tags each row with trendingTopic. Can be used alone or mixed with other sources.

## `trendingMaxResults` (type: `integer`):

Max videos from topic-hub discovery when Include trending is on.

## `scrapeDetails` (type: `boolean`):

Fetch richer metadata for each video, such as exact publish date, category, fuller description, and refreshed view count when available. Best used on smaller runs or shortlisted Video URLs.

## `includeTranscript` (type: `boolean`):

Attempt to fetch timestamped captions and full transcript text for each video. Transcript fields are added only when captions are available and can be fetched. This is the heaviest mode, so it is best for targeted Video URLs or small shortlists.

## `transcriptLanguage` (type: `string`):

Preferred transcript language code, for example en, es, fr, de, ja, or pt. If the requested language is unavailable, the Actor can fall back to the first available caption track.

## Actor input object example

```json
{
  "runMode": "real",
  "workerBaseUrl": "https://yt.opendata.best",
  "searchQueries": [
    "ai workflow automation",
    "youtube competitor analysis"
  ],
  "channelUrls": [
    "https://www.youtube.com/@freecodecamp"
  ],
  "videoUrls": [
    "https://www.youtube.com/watch?v=XVv6mJpFOb0"
  ],
  "maxResults": 50,
  "gl": "US",
  "hl": "en",
  "proxyCountry": "US",
  "uploadDate": "any",
  "resultType": "any",
  "duration": "any",
  "channelTab": "videos",
  "includeChannelAbout": false,
  "useProxyForSearch": false,
  "concurrency": 6,
  "includeComments": false,
  "maxComments": 0,
  "includeCommentReplies": true,
  "maxRepliesPerComment": 0,
  "commentSort": "top",
  "includeRelated": false,
  "maxRelated": 20,
  "includeTrending": false,
  "trendingMaxResults": 50,
  "scrapeDetails": false,
  "includeTranscript": false,
  "transcriptLanguage": "en"
}
```

# Actor output Schema

## `videos` (type: `string`):

Dataset containing one row per unique YouTube video with metadata, optional detail enrichment, and optional transcript fields.

## `runSummary` (type: `string`):

Key-value store record with input echo, source counters, enrichment counters, error counts, and finish timestamp.

## `inputEcho` (type: `string`):

Key-value store record containing the sanitized Actor input used for this run.

## `errorSummary` (type: `string`):

Key-value store record with controlled failure or degraded-run details when present.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "ai workflow automation",
        "youtube competitor analysis"
    ],
    "channelUrls": [
        "https://www.youtube.com/@freecodecamp"
    ],
    "videoUrls": [
        "https://www.youtube.com/watch?v=XVv6mJpFOb0"
    ],
    "gl": "US",
    "hl": "en",
    "transcriptLanguage": "en"
};

// Run the Actor and wait for it to finish
const run = await client.actor("lentic_clockss/youtube-research-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": [
        "ai workflow automation",
        "youtube competitor analysis",
    ],
    "channelUrls": ["https://www.youtube.com/@freecodecamp"],
    "videoUrls": ["https://www.youtube.com/watch?v=XVv6mJpFOb0"],
    "gl": "US",
    "hl": "en",
    "transcriptLanguage": "en",
}

# Run the Actor and wait for it to finish
run = client.actor("lentic_clockss/youtube-research-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "ai workflow automation",
    "youtube competitor analysis"
  ],
  "channelUrls": [
    "https://www.youtube.com/@freecodecamp"
  ],
  "videoUrls": [
    "https://www.youtube.com/watch?v=XVv6mJpFOb0"
  ],
  "gl": "US",
  "hl": "en",
  "transcriptLanguage": "en"
}' |
apify call lentic_clockss/youtube-research-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=lentic_clockss/youtube-research-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/NEhc1Isa4dpa72xIF/builds/iKxodlmKFeU8sGQpw/openapi.json
