# Substack Scraper (`apium/substack-scraper`) Actor

Extract posts, articles, and newsletter data from any Substack publication. Get titles, full text, authors, dates, likes, and comment counts. Export as JSON or CSV, or feed into AI pipelines.

- **URL**: https://apify.com/apium/substack-scraper.md
- **Developed by:** [Tommi Sullivan](https://apify.com/apium) (community)
- **Categories:** News, AI, Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Substack Scraper — extract posts, articles & newsletter data

Substack Scraper lets you **extract posts and articles from any Substack newsletter** in seconds. Give it one or more publication URLs and get structured data: titles, full article text, authors, publish dates, likes, comment counts, cover images, and paywall status — ready to export as JSON, CSV, or Excel, or feed straight into an AI/LLM pipeline.

### What can Substack Scraper do?

- ✅ Scrape **all posts** from any Substack publication (newest, top, or community-picked)
- ✅ Extract **full article text** (clean plain text + original HTML) — perfect for RAG, LLM training data, and content analysis
- ✅ Get engagement data: **likes and comment counts** for every post
- ✅ Detect **paywalled vs. free** posts, with an option to skip subscriber-only content
- ✅ Collect **author names and handles** for every post
- ✅ Scrape multiple newsletters in a single run

### Why scrape Substack?

Substack hosts hundreds of thousands of newsletters covering tech, finance, politics, and culture. Use this scraper for competitor and trend monitoring, building AI knowledge bases from newsletter archives, content research and curation, tracking writer engagement over time, or archiving publications you follow.

### Input

| Field | Description |
|---|---|
| `publicationUrls` | One or more Substack homepages, e.g. `https://example.substack.com` (custom domains work too) |
| `maxPostsPerPublication` | Cap the number of posts per newsletter (default 50) |
| `includeFullText` | Fetch complete article body for every post (default on) |
| `sortBy` | `new`, `top`, or `community` |
| `onlyFreePosts` | Skip paywalled posts |

### Output example

```json
{
  "title": "How I built a $1M side project",
  "url": "https://example.substack.com/p/how-i-built-it",
  "publishedAt": "2026-06-21T14:02:11.000Z",
  "isPaywalled": false,
  "likes": 412,
  "comments": 87,
  "wordCount": 2350,
  "authors": [{ "name": "Jane Doe", "handle": "janedoe" }],
  "bodyText": "Full article text here..."
}
```

### How much does it cost?

You pay a small fee per post scraped (pay-per-result). A typical 50-post newsletter archive costs just a few cents. There are no subscriptions — you only pay for what you extract.

### Using the results

Export to JSON, CSV, Excel, or HTML from the Apify Console, connect via [API](https://docs.apify.com/api/v2), or integrate with Zapier, Make, LangChain, and LlamaIndex. Schedule the Actor to monitor newsletters daily or weekly.

### Is it legal to scrape Substack?

This Actor only collects publicly available data — posts anyone can open in a browser. It does not bypass paywalls or collect private data. Always review the terms of service of the websites you scrape and consult a lawyer if unsure about your specific use case.

### FAQ

**Does it work with custom domains?** Yes — any Substack-powered publication works, e.g. `newsletter.pragmaticengineer.com`.

**Can it bypass paywalls?** No. Paywalled posts return metadata only (title, date, likes), never subscriber-only content.

**Can I run it on a schedule?** Yes, use Apify's built-in scheduler to monitor publications and get only fresh posts.

# Actor input Schema

## `publicationUrls` (type: `array`):

One or more Substack publication homepages, e.g. https://newsletter.pragmaticengineer.com or https://example.substack.com. Post URLs also work — the publication will be detected automatically.

## `maxPostsPerPublication` (type: `integer`):

Maximum number of posts to scrape from each publication (newest first).

## `includeFullText` (type: `boolean`):

Fetch the complete article body (as clean text + HTML) for every post. Slower but gives you the full content — ideal for AI/LLM pipelines.

## `sortBy` (type: `string`):

Order in which posts are collected.

## `onlyFreePosts` (type: `boolean`):

Only return posts that are free to read (skips subscriber-only posts).

## `proxyConfiguration` (type: `object`):

Proxy settings. Datacenter proxies work fine for Substack.

## Actor input object example

```json
{
  "publicationUrls": [
    {
      "url": "https://newsletter.pragmaticengineer.com"
    }
  ],
  "maxPostsPerPublication": 50,
  "includeFullText": true,
  "sortBy": "new",
  "onlyFreePosts": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "publicationUrls": [
        {
            "url": "https://newsletter.pragmaticengineer.com"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("apium/substack-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "publicationUrls": [{ "url": "https://newsletter.pragmaticengineer.com" }] }

# Run the Actor and wait for it to finish
run = client.actor("apium/substack-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "publicationUrls": [
    {
      "url": "https://newsletter.pragmaticengineer.com"
    }
  ]
}' |
apify call apium/substack-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=apium/substack-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/aHWG9vGNLSGzbRFnj/builds/ao9a8sdNOr02xWPFZ/openapi.json
