# Video to Text | TikTok, Reels, Douyin, B站, RedNote Transcript (`ethereal_wool/any-url-transcript`) Actor

Paste any video link — TikTok, Instagram Reels, Douyin, Bilibili, Xiaohongshu (RedNote) — and get the spoken words back as text. Real AI speech recognition, no captions needed. Full transcript + timestamped sentences as clean JSON.

- **URL**: https://apify.com/ethereal\_wool/any-url-transcript.md
- **Developed by:** [Jackie Chen](https://apify.com/ethereal_wool) (community)
- **Categories:** Social media, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$25.00 / 1,000 transcript minutes

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Any URL Transcript — Video to Text for TikTok, Reels, Douyin, B站 & RedNote

**One Actor, every short-video platform.** Paste video links from **TikTok,
Instagram Reels, Douyin (抖音), Bilibili (B站), or Xiaohongshu (小红书 /
RedNote)** — in any mix — and get back the **full spoken transcript plus
timestamped sentences** as clean JSON. Real AI speech recognition, ready for
LLMs, RAG pipelines, content research, and subtitle workflows.

> **This is real ASR, not caption scraping.** Most "transcript" tools only
> download captions a creator happened to upload — and most short videos have
> none. This Actor runs industrial speech recognition on each video's audio
> track, so it works on **any video with speech**, captions or not. Chinese
> (Mandarin) recognition is best-in-class; English, Cantonese, Japanese, and
> Korean are supported too.

> **Unofficial.** This Actor is not affiliated with, authorized, or endorsed by
> TikTok, ByteDance, Meta, Bilibili, or Xiaohongshu. It is an independent tool
> that processes publicly available content. Use it in compliance with each
> platform's terms and all applicable laws; you are responsible for how you use
> the retrieved data.

### Supported platforms

| Platform | Accepted links |
|---|---|
| **TikTok** | `tiktok.com/@user/video/…`, `vm.tiktok.com/…`, `vt.tiktok.com/…` |
| **Douyin 抖音** | `douyin.com/video/…`, `v.douyin.com/…`, pasted share text |
| **Instagram Reels** | `instagram.com/reel/…`, `/p/…`, `/tv/…` |
| **Bilibili B站** | `bilibili.com/video/BV…`, `b23.tv/…`, bare BV ids |
| **Xiaohongshu 小红书** | `xiaohongshu.com/explore/…`, `xhslink.com/…` |

The Actor auto-detects the platform per URL — mix them freely in one run.

### What you get per video

- **Full transcript** (`fullText`) — the complete spoken content as one string.
- **Timestamped sentences** (`sentences[]`) — each with `startMs` / `endMs`,
  ready for subtitles, deep links, and clip selection.
- **Platform + metadata** — `platform`, author, title/caption, duration,
  publish time, play / like / comment / share counts where the platform
  provides them.

### Quick start

1. Open the Actor and press **Run** — the default input works out of the box.
2. Replace the examples with your own links from any supported platform.
3. Language defaults to **Auto-detect**; set it explicitly for single-language
   batches to improve accuracy (Chinese is best-in-class).
4. Collect results from the **Dataset** tab as JSON / CSV / Excel, or pull them
   via the [Apify API](https://docs.apify.com/api/v2) and MCP from your own code.

No proxies, no cookies, no login — everything runs server-side.

### Example output

```json
{
  "platform": "tiktok",
  "videoId": "7405073368414408965",
  "videoUrl": "https://www.tiktok.com/@khaby.lame/video/7405073368414408965",
  "author": "khaby.lame",
  "durationSec": 18.2,
  "playCount": 24500000,
  "likeCount": 2900000,
  "language": "auto",
  "detectedSpeech": true,
  "sentenceCount": 4,
  "fullText": "You don't need an expensive gadget for this…",
  "sentences": [
    { "text": "You don't need an expensive gadget for this.", "startMs": 410, "endMs": 2750 }
  ]
}
```

### What people build with it

- **Cross-platform content research** — compare what creators say about the
  same topic on TikTok, Reels, and the Chinese platforms, in one dataset.
- **Viral-hook mining** — transcribe the winners across platforms and study
  the exact opening lines that earn views everywhere.
- **Agent & workflow integration** — one endpoint for "video URL → text" means
  your agent doesn't need to care which platform a link came from.
- **LLM & RAG pipelines** — build multilingual short-video speech corpora with
  engagement scores attached.
- **Subtitles & translation** — timestamped sentences drop straight into
  SRT/VTT generation and dubbing workflows.

### Pricing & billing

Pay per started transcript minute (**$0.025/minute**, one-minute minimum) on
every supported platform. Failed fetches, deleted videos, image posts,
unsupported URLs, videos over your duration cap, and failed transcriptions are
**not charged**.

### Why this Actor

- **One input for five platforms** — no more wiring a separate scraper per
  platform into your pipeline.
- **Real speech recognition** — works without captions; Chinese (Mandarin)
  accuracy is best-in-class, where Western tools struggle most.
- **Cheapest stream selection** — smallest encode per platform, audio-only on
  Bilibili, so runs stay fast.
- **Direct API, no headless browser** — fast, stable runs with nothing to babysit.
- **No login, no cookies** — we never touch your accounts, so there's no ban risk.
- **Structured JSON** — export to CSV, Excel, or JSON, or pull straight from
  the API / MCP.

### Tips for better results

- Feed it the winners: find the top-performing videos in your niche first
  (search / profile scrapers), then transcribe just those.
- Single-language batch? Set the language explicitly — it noticeably improves
  accuracy over Auto.
- `detectedSpeech: false` flags music-only / no-speech videos so you can filter
  them out downstream.
- Long Bilibili videos: raise `maxDurationMinutes` (up to 240) for lectures —
  audio-only downloads keep even those runs quick.

### FAQ

**Do I need an account, cookies, or to log in anywhere?**
No. The Actor talks to fast, direct HTTP APIs server-side — you just provide
video links and run it.

**Does it work on videos without captions?**
Yes — that's the point. It runs real speech recognition on the audio, so
captions are never required.

**Which languages are supported?**
Auto-detect, Chinese (best-in-class), Cantonese, English, Japanese, and Korean.

**How am I billed?**
$0.025 per started transcript minute on every platform. Videos that can't be
fetched or transcribed are not charged.

**Can I run it on a schedule or call it from my app?**
Yes — use Apify Schedules, the REST API, the JavaScript / Python clients, or
the MCP server. See the **API** tab.

**Is this affiliated with the platforms?**
No. It's an independent tool that processes publicly available content. Use it
in line with each platform's terms and applicable law.

# Actor input Schema

## `videoUrls` (type: `array`):

Video links to transcribe, in any mix of platforms: TikTok, Douyin (抖音), Instagram Reels, Bilibili (B站), Xiaohongshu (小红书 / RedNote). Short share links (vm.tiktok.com, v.douyin.com, b23.tv, xhslink.com) work too. One transcript is produced per video.

## `language` (type: `string`):

Language spoken in the videos. Auto-detect works well for mixed batches; setting an explicit language noticeably improves accuracy for single-language content (Chinese accuracy is best-in-class).

## `maxDurationMinutes` (type: `integer`):

Videos longer than this are skipped (and not charged). Mostly relevant for Bilibili, which hosts multi-hour videos. Raise it up to 240 for long lectures.

## `proxyConfiguration` (type: `object`):

Optional. Route the upstream API calls through an Apify Proxy to vary the source IP. Usually not needed.

## Actor input object example

```json
{
  "videoUrls": [
    "https://www.tiktok.com/@khaby.lame/video/7405073368414408965",
    "https://www.bilibili.com/video/BV1HjmrBzEDA/"
  ],
  "language": "auto",
  "maxDurationMinutes": 60,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoUrls": [
        "https://www.tiktok.com/@khaby.lame/video/7405073368414408965",
        "https://www.bilibili.com/video/BV1HjmrBzEDA/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ethereal_wool/any-url-transcript").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videoUrls": [
        "https://www.tiktok.com/@khaby.lame/video/7405073368414408965",
        "https://www.bilibili.com/video/BV1HjmrBzEDA/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("ethereal_wool/any-url-transcript").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoUrls": [
    "https://www.tiktok.com/@khaby.lame/video/7405073368414408965",
    "https://www.bilibili.com/video/BV1HjmrBzEDA/"
  ]
}' |
apify call ethereal_wool/any-url-transcript --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=ethereal_wool/any-url-transcript",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/RgcnX9WU6Gfgf2STc/builds/TQaeF7Cb8xAcjhQQV/openapi.json
