# Bilibili Transcript Scraper | AI Speech-to-Text (B站) (`ethereal_wool/bilibili-transcript-scraper`) Actor

Turn any Bilibili (B站) video into text. Real AI speech recognition with best-in-class Mandarin Chinese accuracy — works on videos with no subtitles. Full text + timestamped sentences as clean JSON.

- **URL**: https://apify.com/ethereal\_wool/bilibili-transcript-scraper.md
- **Developed by:** [Jackie Chen](https://apify.com/ethereal_wool) (community)
- **Categories:** Social media, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$25.00 / 1,000 transcript minutes

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Bilibili Transcript Scraper — AI Speech-to-Text for B站 Videos

Turn any **Bilibili (B站) video into text** with real AI speech recognition.
Paste video links, get back the **full spoken transcript plus timestamped
sentences** as clean JSON — with **best-in-class Mandarin Chinese accuracy**,
ready for LLMs, RAG pipelines, content research, and subtitle workflows.

> **This is real ASR, not subtitle scraping.** Most "transcript" tools only
> download subtitles a creator happened to upload — and the majority of
> Bilibili videos have none. This Actor runs industrial speech recognition on
> the video's audio track, so it works on **any Bilibili video with speech**,
> subtitles or not. Chinese (Mandarin) recognition is its strongest language.

> **Unofficial.** This Actor is not affiliated with, authorized, or endorsed by
> Bilibili. It is an independent tool that processes publicly available
> content. Use it in compliance with Bilibili's terms and all applicable laws;
> you are responsible for how you use the retrieved data.

### What you get per video

- **Full transcript** (`fullText`) — the complete spoken content as one string.
- **Timestamped sentences** (`sentences[]`) — each with `startMs` / `endMs`,
  ready for subtitles, deep links, and clip selection.
- **Video metadata** — title, UP主 (author), duration, publish time, view /
  like / coin / favorite / danmaku / comment / share counts, and cover image
  URL, so every transcript arrives with its engagement context attached.

### Quick start

1. Open the Actor and press **Run** — the default input works out of the box.
2. Replace the example with your own video links: full URLs
   (`https://www.bilibili.com/video/BV…`), short share links
   (`https://b23.tv/XXXX`), or bare BV ids.
3. Language defaults to **Chinese**; switch to Auto for mixed or non-Chinese
   content.
4. Collect results from the **Dataset** tab as JSON / CSV / Excel, or pull them
   via the [Apify API](https://docs.apify.com/api/v2) and MCP from your own code.

No proxies, no cookies, no login — everything runs server-side.

### Example output

```json
{
  "videoId": "BV1HjmrBzEDA",
  "videoUrl": "https://www.bilibili.com/video/BV1HjmrBzEDA/",
  "title": "Ultra级旗舰!OPPO Find X9 Pro好用到超标 #数码生活 #科技数码",
  "author": "科技小辛",
  "durationSec": 87,
  "viewCount": 152000,
  "likeCount": 8400,
  "coinCount": 1200,
  "danmakuCount": 430,
  "language": "zh",
  "detectedSpeech": true,
  "sentenceCount": 13,
  "fullText": "如果有一台手机能让果粉变墨粉，那一定是OPPO Find X9 Pro…",
  "sentences": [
    { "text": "如果有一台手机能让果粉变墨粉，", "startMs": 280, "endMs": 2350 },
    { "text": "那一定是OPPO Find X9 Pro。", "startMs": 2350, "endMs": 4400 }
  ]
}
```

### What people build with it

- **Chinese-market research** — index what B站 UP主 actually *say* (not just
  titles) to understand trends, opinions, and product sentiment in your category.
- **Long-form knowledge mining** — Bilibili is China's home of tech reviews,
  tutorials, and video essays; transcripts turn hours of video into searchable,
  quotable text.
- **Viral-hook mining** — pull transcripts of the top videos in your niche and
  study the exact opening lines and structures that earn views and coins.
- **Cross-border content** — transcribe Chinese videos, then translate and
  repurpose them for YouTube or your own market.
- **LLM & RAG pipelines** — build Chinese-language corpora from real video
  speech, with engagement scores as a free quality signal.
- **Subtitles & translation** — timestamped sentences drop straight into
  SRT/VTT generation and dubbing workflows.

### Pricing & billing

Flat **pay-per-transcript**: you're charged once per successfully transcribed
video. Failed fetches, deleted or region-locked videos, videos over your
duration cap, and failed transcriptions are **not charged**. No separate
compute or platform fee — the price you see is the price you pay.

### Why this Actor

- **Best-in-class Chinese ASR** — Mandarin recognition is its strongest
  language, where Western transcript tools struggle most.
- **Works without subtitles** — real speech recognition on the audio track;
  most Bilibili videos carry no CC subtitles to scrape.
- **Audio-only downloads** — Bilibili serves audio as a separate stream, so the
  Actor fetches megabytes instead of the full video. Long videos transcribe
  fast and cheap.
- **Engagement context included** — every transcript ships with view / like /
  coin / favorite / danmaku counts, so you can rank by performance immediately.
- **Direct API, no headless browser** — fast, stable runs with nothing to babysit.
- **No login, no cookies** — we never touch your accounts, so there's no ban risk.
- **Structured JSON** — export to CSV, Excel, or JSON, or pull straight from
  the API / MCP.

### Tips for better results

- Feed it the winners: find the top-performing videos in your niche first
  (search / channel scrapers), then transcribe just those.
- Keep the language on **Chinese** for B站 content — it noticeably improves
  accuracy over Auto.
- The default duration cap is 60 minutes; raise `maxDurationMinutes` (up to
  240\) for lectures and documentaries, or lower it to keep batch runs snappy.
- For multi-part videos (分P), part 1 is transcribed; submit other parts'
  direct URLs if you need them.
- `detectedSpeech: false` flags music-only / no-speech videos so you can filter
  them out downstream.

### FAQ

**Do I need an account, cookies, or to log in anywhere?**
No. The Actor talks to fast, direct HTTP APIs server-side — you just provide
video links and run it.

**Does it work on videos without subtitles?**
Yes — that's the point. It runs real speech recognition on the audio, so
subtitles are never required (and most Bilibili videos don't have them).

**How good is the Chinese accuracy?**
Mandarin is the model's strongest language; it handles fast colloquial speech,
regional accents, and technical vocabulary well, and returns punctuated
sentences.

**What about really long videos?**
Videos longer than your `maxDurationMinutes` cap (default 60) are skipped and
not charged. Raise the cap up to 240 minutes when you need lectures or
documentaries — audio-only downloads keep even those runs quick.

**How am I billed?**
One fixed price per successfully transcribed video. Videos that can't be
fetched or transcribed are not charged.

**Can I run it on a schedule or call it from my app?**
Yes — use Apify Schedules, the REST API, the JavaScript / Python clients, or
the MCP server. See the **API** tab.

**Is this affiliated with Bilibili?**
No. It's an independent tool that processes publicly available content. Use it
in line with the platform's terms and applicable law.

# Actor input Schema

## `videoUrls` (type: `array`):

Bilibili video links to transcribe. Accepts full URLs (`https://www.bilibili.com/video/BV…`), short share links (`https://b23.tv/XXXX`), or bare BV ids. One transcript is produced per video (part 1 for multi-part videos).

## `language` (type: `string`):

Language spoken in the videos. Defaults to Chinese (B站 is Mandarin-heavy and Chinese accuracy is best-in-class). Choose Cantonese for 粤语 content, or Auto for mixed / non-Chinese audio.

## `maxDurationMinutes` (type: `integer`):

Videos longer than this are skipped (and not charged). Bilibili hosts multi-hour videos; the default keeps runs fast and predictable. Raise it up to 240 for long lectures or documentaries.

## `proxyConfiguration` (type: `object`):

Optional. Route the upstream API calls through an Apify Proxy to vary the source IP. Usually not needed.

## Actor input object example

```json
{
  "videoUrls": [
    "https://www.bilibili.com/video/BV1HjmrBzEDA/"
  ],
  "language": "zh",
  "maxDurationMinutes": 60,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoUrls": [
        "https://www.bilibili.com/video/BV1HjmrBzEDA/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ethereal_wool/bilibili-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videoUrls": ["https://www.bilibili.com/video/BV1HjmrBzEDA/"] }

# Run the Actor and wait for it to finish
run = client.actor("ethereal_wool/bilibili-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoUrls": [
    "https://www.bilibili.com/video/BV1HjmrBzEDA/"
  ]
}' |
apify call ethereal_wool/bilibili-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=ethereal_wool/bilibili-transcript-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/RJedgDeWEZgarHyXw/builds/ktf4X2x2igFMet5oT/openapi.json
