# Zhihu \[Just 💰$3] — Hot List, Q\&A & Author Profiles (`blackfalcondata/zhihu-scraper`) Actor

💰 $3 per 1,000 items. Scrape zhihu.com — trending hot-list questions (热榜), full Q\&A answers with text and engagement counts, and author profiles as structured data. No login or API key required. Incremental mode flags new and changed records for monitoring and AI pipelines.

- **URL**: https://apify.com/blackfalcondata/zhihu-scraper.md
- **Developed by:** [Black Falcon Data](https://apify.com/blackfalcondata) (community)
- **Categories:** Social media, Lead generation, Automation
- **Stats:** 3 total users, 3 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

### What does Zhihu do?

Zhihu Scraper extracts structured Q\&A data from [zhihu.com](https://zhihu.com) — trending hot-list questions, full question and answer text, and author profiles. Each record carries engagement metrics (upvotes, comments, thanks, and followers), with no login required.

### How to use this actor

- 👉 **Register for a free Apify account** — no credit card required.
- 🎉 Just click **[Sign up free on Apify →](https://console.apify.com/sign-up?fpr=1h3gvi\&fp_sid=ctarich)** and complete a quick signup.
- 💰 A free Apify account includes $5 in monthly credits — enough to test this actor.
- ⏳ Scrape during the free trial, with no commitment or upfront payment required.

### Key features

- **🔔 Notifications** — Telegram, Slack, Discord, WhatsApp Cloud API, and generic webhook out of the box. Pair with incremental for daily new-listing alerts without pipeline glue.
- **🔗 Paste-mode** — paste any zhihu URL straight from your browser — single-listing pages, search-results URLs, or category SEO URLs. Mix freely with keyword and IDs in the same run; results dedupe by ID.
- **📧 Email + phone extraction** — best-effort regex extraction of contact emails and phone numbers from descriptions — emitted as `extractedEmails[]` and `extractedPhones[]` on every record.
- **🔗 URL + social-profile extraction** — every record carries `extractedUrls[]` plus a structured `socialProfiles { linkedin, twitter, instagram, facebook, youtube, tiktok, github, xing }` parsed from the description.
- **📦 Compact mode** — AI-agent and MCP-friendly payloads with core fields only.
- **📌 Change classification** — each record carries a `changeType` of NEW / UPDATED / UNCHANGED / REAPPEARED / EXPIRED. Default emits NEW + UPDATED + REAPPEARED; opt into the others with `emitUnchanged` / `emitExpired`. Repost detection flags previously-expired listings that come back.
- **🔌 MCP connectors** — export your results into Notion via Apify's MCP connectors — a clean run-summary page, no glue code. Opt-in via the App connector field; deterministic field-mapping, no AI. Built on Apify's connector framework, so more destinations open up as their catalog grows.
- **♻️ Incremental mode** — recurring runs emit and charge only for listings that are new or whose tracked content changed. First run builds the baseline; subsequent runs emit only NEW / UPDATED / REAPPEARED records (UNCHANGED + EXPIRED opt-in). Saves 80–95% on daily monitoring.
- **🧹 Empty-field stripping** — drop null, empty-string, and empty-array fields from each record before push. Smaller payloads for AI agents and dashboards that already handle missing fields gracefully.
- **📤 Export anywhere** — Download the dataset as JSON, CSV, or Excel from the Apify Console, or stream live via the Apify API and integrations (Make, Zapier, Google Sheets, n8n, …).

### What data can you extract from zhihu.com?

Each result includes Core listing fields (`type`, `id`, `recordId`, `url`, `title`, `excerpt`, `content`, and `contentLength`, and more). In standard mode, all fields are always present — unavailable data points are returned as `null`, never omitted. In compact mode, only core fields are returned.

Enable detail enrichment in the input to get richer fields such as full descriptions where the source provides them.

### Input

The main inputs are a result limit. Additional filters and options are available in the input schema.

Key parameters:

- **`operation`** — What to scrape. **Hot List** = today's trending questions (热榜). **Questions** = full question + answers for the URLs you supply. **Profiles** = author profile stats for the URLs you supply. (default: `"hotList"`)
- **`limit`** — Number of trending questions to fetch (Hot List operation). 1–100. (default: `50`)
- **`startUrls`** — For **Questions**: question URLs (https://www.zhihu.com/question/12345) or bare ids. For **Profiles**: profile URLs (https://www.zhihu.com/people/token) or bare tokens.
- **`includeAnswers`** — Fetch each question's page for the full question detail + first-page answers (with full text and engagement). Turn off for a fast trending snapshot (Hot List metadata only). (default: `true`)
- **`maxAnswersPerQuestion`** — Cap answers emitted per question. 0 = all embedded first-page answers. (default: `0`)
- **`compact`** — Output only core fields (for AI-agent / MCP workflows). (default: `false`)
- **`excludeEmptyFields`** — Drop null, empty-string, and empty-array fields from each record before push. (default: `false`)
- **`incrementalMode`** — Compare against previous run state and tag each record NEW / CHANGED / UNCHANGED. stateKey is optional — defaults to a stable key derived from the operation and targets. (default: `false`)
- **`stateKey`** — Optional. Stable identifier for the tracked set (e.g. "zhihu-hotlist"). Leave empty to auto-generate.
- **`emitUnchanged`** — When incremental, also emit records that haven't changed. (default: `false`)
- **`telegramToken`** — Telegram bot token (from @BotFather). Required for Telegram notifications.
- **`telegramChatId`** — Telegram chat or channel ID (e.g. "-100123456789"). Required when telegramToken is set.
- ...and 16 more parameters

### Input examples

**Trending hot list (fast snapshot)** — undefined

→ undefined

```json
{
  "operation": "hotList",
  "limit": 20,
  "includeAnswers": false
}
```

**Hot list with full answers** — undefined

→ undefined

```json
{
  "operation": "hotList",
  "limit": 10,
  "includeAnswers": true
}
```

**Specific questions by URL** — undefined

→ undefined

```json
{
  "operation": "questions",
  "startUrls": [
    "https://www.zhihu.com/question/19550225"
  ]
}
```

**Author profiles by URL** — undefined

→ undefined

```json
{
  "operation": "profiles",
  "startUrls": [
    "https://www.zhihu.com/people/zhang-jia-wei"
  ]
}
```

**Track the hot list for changes (incremental)** — undefined

→ undefined

```json
{
  "operation": "hotList",
  "limit": 50,
  "includeAnswers": false,
  "incrementalMode": true
}
```

### Output

Each run produces a dataset of structured listing records. Results can be downloaded as JSON, CSV, or Excel from the Dataset tab in Apify Console.

### Example listing record

```json
{
  "type": "question",
  "id": "2061371759589323800",
  "recordId": "zhihu.com:question:2061371759589323800",
  "url": "https://www.zhihu.com/question/2061371759589323800",
  "title": "经济学家任泽平 VIP 付费会员群「暴雷」，有人听信操作建议亏损 1000多万，你如何看这种现象？",
  "excerpt": "科技股持续下跌，网红经济学家任泽平站上风口浪尖。此前，任泽平持续看好科技牛行情，指出AI科技牛是康波周期量级的机会，天花板远远没有看到，是这代人最重要的时代机遇。 此前科技股走出了波澜壮阔的行情，双创指数均创下历史新高，然而近期韩国“存储双雄”跳水，引发A股连锁反应，科技股遭遇踩踏行情。 据报道，近期群名为“泽平宏观VIP群30”的付费会员群出现投资者激烈控诉。该投资者称，自己听信任泽平相关 “科...",
  "contentLength": 0,
  "commentCount": 0,
  "answerCount": 192,
  "followerCount": 439,
  "hotRank": 1,
  "hotHeat": "442 万热度",
  "authorName": "用户",
  "questionId": "2061371759589323800",
  "questionTitle": "经济学家任泽平 VIP 付费会员群「暴雷」，有人听信操作建议亏损 1000多万，你如何看这种现象？",
  "topics": [
    "922",
    "6722",
    "6723",
    "19800",
    "166889"
  ],
  "createdTime": "2026-07-17T00:48:45.000Z",
  "fetchedAt": "2026-07-18T00:00:00.000Z"
}
```

### Incremental fields

When incremental mode is on, each record also carries:

- `changeType` — one of `NEW`, `UPDATED`, `UNCHANGED`, `REAPPEARED`, `EXPIRED`. Default output covers `NEW` / `UPDATED` / `REAPPEARED`; set `emitUnchanged: true` or `emitExpired: true` to opt into the others.
- `isRepost`, `repostOfId`, `repostDetectedAt` — populated when a new listing matches the tracked content of a previously expired one. Set `skipReposts: true` to drop detected reposts from the output.

### How to scrape zhihu.com

1. Go to [Zhihu](https://apify.com/blackfalcondata/zhihu-scraper?fpr=1h3gvi) in Apify Console.
2. Configure the input.
3. Set `maxResults` to control how many results you need.
4. Enable `includeDetails` if you need full descriptions.
5. Click **Start** and wait for the run to finish.
6. Export the dataset as JSON, CSV, or Excel.

### Use cases

- Extract listing data from zhihu.com for market research and competitive analysis.
- Monitor new and changed listings on scheduled runs without processing the full dataset every time.
- Feed structured data into AI agents, MCP tools, and automated pipelines using compact mode.
- Export clean, structured data to dashboards, spreadsheets, or data warehouses.

### How much does it cost to scrape zhihu.com?

Zhihu uses [pay-per-event](https://docs.apify.com/platform/actors/paid-actors/pay-per-event) pricing. You pay a small fee when the run starts and then for each result that is actually produced.

- **Run start:** $0.01 per run
- **Per result:** $0.003 per listing record

Example costs:

- 10 results: **$0.04**
- 25 results: **$0.085**
- 100 results: **$0.31**
- 200 results: **$0.61**
- 500 results: **$1.51**

#### Example: recurring monitoring savings

These examples compare full re-scrapes with incremental runs at different churn rates. Churn is the share of listings that are new or whose tracked content changed since the previous run. Actual churn depends on your query breadth, source activity, and polling frequency — the scenarios below are examples, not predictions.

Example setup: 100 listings per run, daily polling (30 runs/month). Costs scale linearly with the number of listings.

| Churn rate | Full re-scrape run cost | Incremental run cost | Savings vs full re-scrape | Monthly cost after baseline |
|---|---:|---:|---:|---:|
| 5% — stable niche query | $0.31 | $0.03 | $0.28 (92%) | $0.75 |
| 15% — moderate broad query | $0.31 | $0.06 | $0.26 (82%) | $1.65 |
| 30% — high-volume aggregator | $0.31 | $0.10 | $0.21 (68%) | $3.00 |

Full re-scrape monthly cost at the same cadence: $9.30. First month with incremental costs $1.04 / $1.91 / $3.21 for the 5% / 15% / 30% scenarios because the first run builds baseline state at full cost before incremental savings apply.

Platform usage (compute and proxies) is billed separately by Apify based on actual consumption. Incremental runs consume less on result processing, though fixed per-run overhead stays the same.

### FAQ

#### How many results can I get from zhihu.com?

The number of results depends on the search query and available listings on zhihu.com. Use the `maxResults` parameter to control how many results are returned per run.

#### Does Zhihu support recurring monitoring?

Yes. Enable incremental mode to only receive new or changed listings on subsequent runs. This is ideal for scheduled monitoring where you want to track changes over time without re-processing the full dataset.

#### Can I integrate Zhihu with other apps?

Yes. Zhihu works with Apify's [integrations](https://apify.com/integrations?fpr=1h3gvi) to connect with tools like Zapier, Make, Google Sheets, Slack, and more. You can also use webhooks to trigger actions when a run completes.

#### Can I use Zhihu with the Apify API?

Yes. You can start runs, manage inputs, and retrieve results programmatically through the [Apify API](https://docs.apify.com/api/v2). Client libraries are available for JavaScript, Python, and other languages.

#### Can I use Zhihu through an MCP Server?

Yes. Apify provides an [MCP Server](https://apify.com/apify/actors-mcp-server?fpr=1h3gvi) that lets AI assistants and agents call this actor directly. Use compact mode, a single `descriptionFormat`, and `excludeEmptyFields` to keep payloads manageable for LLM context windows.

#### Is it legal to scrape zhihu.com?

This actor extracts publicly available data from zhihu.com. Web scraping of public information is generally considered legal, but you should always review the target site's terms of service and ensure your use case complies with applicable laws and regulations, including GDPR where relevant.

#### Your feedback

If you have questions, need a feature, or found a bug, please [open an issue](https://apify.com/blackfalcondata/zhihu-scraper/issues?fpr=1h3gvi) on the actor's page in Apify Console. Your feedback helps us improve.

### You might also like

- [Douyin \[Just 💰$0.5\] — Hot Search, Trending & Viral](https://apify.com/blackfalcondata/douyin-scraper?fpr=1h3gvi) — 💰 $0.50 per 1,000 results. Scrape douyin.com real-time trending boards — hot search, seeding &.
- [Facebook Group Posts — group content & engagement](https://apify.com/blackfalcondata/facebook-group-post-scraper?fpr=1h3gvi) — Scrape facebook.com — post text & timestamps · reaction/comment/share counts · media & optional.
- [LinkedIn Profile Scraper + Email](https://apify.com/blackfalcondata/linkedin-profile-scraper?fpr=1h3gvi) — Scrape LinkedIn profiles by URL, handle, or people-search filters for recruiting research, CRM.
- [Maigret Username OSINT — Verified Account Finder](https://apify.com/blackfalcondata/maigret-username-osint-scraper?fpr=1h3gvi) — Find where a username or handle exists across hundreds of social media networks and websites. Every.
- [Pinterest Scraper — Pins, Boards, Profiles & Engagement](https://apify.com/blackfalcondata/pinterest-scraper?fpr=1h3gvi) — \[💰$1.35/1K] Scrape Pinterest pins, boards and profiles by keyword or Start URL. Export images,.
- [Quora Scraper — Q\&A, Profiles & Spaces](https://apify.com/blackfalcondata/quora-scraper?fpr=1h3gvi) — \[💰$0.8/1K] Scrape Quora questions, answers, profiles, posts, and spaces by search or URL —.
- [Reddit Email Scraper — Emails from Posts & Comments](https://apify.com/blackfalcondata/reddit-email-scraper?fpr=1h3gvi) — Scrape reddit.com — extract email addresses and contact details from posts · comments & user.
- [Reddit Lead Scraper \[Just 💰$2\] — Emails & Socials](https://apify.com/blackfalcondata/reddit-lead-scraper?fpr=1h3gvi) — 💰 $2 per 1,000 leads. Scrape reddit.com — turn any subreddit or keyword into a B2B lead list ·.

### Getting started with Apify

New to Apify? [Create a free account with $5 credit](https://console.apify.com/sign-up?fpr=1h3gvi\&fp_sid=ctarich) — no credit card required.

1. Sign up — $5 platform credit included
2. Open this actor and configure your input
3. Click **Start** — export results as JSON, CSV, or Excel

Need more later? [See Apify pricing](https://apify.com/pricing?fpr=1h3gvi).

### Disclaimer

This actor accesses only publicly available data on zhihu.com. You are responsible for how you use the extracted data — in particular any personal information such as names, phone numbers, or email addresses — and for complying with Zhihu's terms of use, applicable data-protection law (including the GDPR where it applies), and the anti-spam rules of your jurisdiction.

This actor is not affiliated with, endorsed by, or connected to Zhihu.

### Search keywords

zhihu scraper, zhihu api, apify zhihu, zhihu data extraction, zhihu.com scraper, zhihu.com data, zhihu.com api.

# Actor input Schema

## `operation` (type: `string`):

What to scrape. **Hot List** = today's trending questions (热榜). **Questions** = full question + answers for the URLs you supply. **Profiles** = author profile stats for the URLs you supply.

## `limit` (type: `integer`):

Number of trending questions to fetch (Hot List operation). 1–100.

## `startUrls` (type: `array`):

For **Questions**: question URLs (https://www.zhihu.com/question/12345) or bare ids. For **Profiles**: profile URLs (https://www.zhihu.com/people/token) or bare tokens.

## `includeAnswers` (type: `boolean`):

Fetch each question's page for the full question detail + first-page answers (with full text and engagement). Turn off for a fast trending snapshot (Hot List metadata only).

## `maxAnswersPerQuestion` (type: `integer`):

Cap answers emitted per question. 0 = all embedded first-page answers.

## `proxyConfiguration` (type: `object`):

Proxy configuration. Recommended for reliable access.

## `compact` (type: `boolean`):

Output only core fields (for AI-agent / MCP workflows).

## `excludeEmptyFields` (type: `boolean`):

Drop null, empty-string, and empty-array fields from each record before push.

## `incrementalMode` (type: `boolean`):

Compare against previous run state and tag each record NEW / CHANGED / UNCHANGED. stateKey is optional — defaults to a stable key derived from the operation and targets.

## `stateKey` (type: `string`):

Optional. Stable identifier for the tracked set (e.g. "zhihu-hotlist"). Leave empty to auto-generate.

## `emitUnchanged` (type: `boolean`):

When incremental, also emit records that haven't changed.

## `telegramToken` (type: `string`):

Telegram bot token (from @BotFather). Required for Telegram notifications.

## `telegramChatId` (type: `string`):

Telegram chat or channel ID (e.g. "-100123456789"). Required when telegramToken is set.

## `discordWebhookUrl` (type: `string`):

Discord incoming webhook URL. Server Settings → Integrations → Webhooks.

## `slackWebhookUrl` (type: `string`):

Slack incoming webhook URL. api.slack.com/messaging/webhooks.

## `webhookUrl` (type: `string`):

Receives a JSON POST with {metadata, items} after each run. For n8n / Make / Zapier / custom backends.

## `webhookHeaders` (type: `object`):

Optional JSON object of custom headers (e.g. {"Authorization":"Bearer ..."}).

## `notificationLimit` (type: `integer`):

Maximum number of records included in each notification message (1–20).

## `notifyOnlyChanges` (type: `boolean`):

When Incremental Mode is on, only notify for NEW and CHANGED records.

## `appConnector` (type: `string`):

Optional. Pick a connected app under Settings → API & Integrations to receive your results. Best-effort across MCP connectors as Apify expands its catalog.

## `mcpIssueTeam` (type: `string`):

Only when the connected app is an issue tracker: the team the summary issue is created under.

## `whatsappAccessToken` (type: `string`):

WhatsApp Cloud API access token from Meta Business Manager. The recipient must have messaged your business number within the last 24 hours.

## `whatsappPhoneNumberId` (type: `string`):

WhatsApp Business phone number ID (numeric, from Meta dashboard). Required when whatsappAccessToken is set.

## `whatsappTo` (type: `string`):

Recipient phone number in E.164 format without + (e.g. "33612345678"). Recipient must have messaged your business number within the last 24 hours.

## `descriptionFormat` (type: `string`):

Choose which representation of the listing description to include. `all` keeps every variant; the others keep only the selected one.

## `maxResults` (type: `integer`):

Maximum number of listings to return per run. Set to 0 for unlimited.

## `includeDetails` (type: `boolean`):

Fetch the detail page for each listing to retrieve the full description text and contact information. Increases run time and cost. Leave off for a fast, low-cost run.

## `emitExpired` (type: `boolean`):

When incremental mode is on, also emit listings that were seen before but are no longer found.

## `skipReposts` (type: `boolean`):

In incremental mode, skip listings that appear to be reposts of a previously-seen expired listing with matching details.

## Actor input object example

```json
{
  "operation": "hotList",
  "limit": 20,
  "includeAnswers": true,
  "maxAnswersPerQuestion": 0,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "compact": false,
  "excludeEmptyFields": false,
  "incrementalMode": false,
  "emitUnchanged": false,
  "notificationLimit": 5,
  "notifyOnlyChanges": false,
  "descriptionFormat": "all",
  "maxResults": 5,
  "includeDetails": false,
  "emitExpired": false,
  "skipReposts": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "operation": "hotList",
    "limit": 20,
    "includeAnswers": false,
    "excludeEmptyFields": false,
    "descriptionFormat": "all",
    "maxResults": 5,
    "includeDetails": false
};

// Run the Actor and wait for it to finish
const run = await client.actor("blackfalcondata/zhihu-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "operation": "hotList",
    "limit": 20,
    "includeAnswers": False,
    "excludeEmptyFields": False,
    "descriptionFormat": "all",
    "maxResults": 5,
    "includeDetails": False,
}

# Run the Actor and wait for it to finish
run = client.actor("blackfalcondata/zhihu-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "operation": "hotList",
  "limit": 20,
  "includeAnswers": false,
  "excludeEmptyFields": false,
  "descriptionFormat": "all",
  "maxResults": 5,
  "includeDetails": false
}' |
apify call blackfalcondata/zhihu-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=blackfalcondata/zhihu-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cR6gMhDxDTmuB2b8r/builds/Nua5unnkGmNuWQ8f3/openapi.json
