# Podcast Email Scraper - Host & Contact Finder (`ritaya/podcast-email-scraper`) Actor

Find podcasts by topic and extract the host/business email, author, website and show metadata from the podcast RSS feed. No login.

- **URL**: https://apify.com/ritaya/podcast-email-scraper.md
- **Developed by:** [Zane](https://apify.com/ritaya) (community)
- **Categories:** Lead generation, Social media
- **Stats:** 3 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $30.00 / 1,000 podcast scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Podcast Email Scraper - Host & Contact Finder

**Podcast email scraper** and **host contact finder**: discover podcasts by topic, then
extract the **host / business email**, author, website, genre and show metadata straight
from each show's RSS feed. Built for **podcast guesting, sponsorship outreach, PR and
B2B-to-creator lead generation**.

Keywords: podcast email scraper, podcast host email, podcast contact finder, podcast
outreach leads, sponsorship prospecting, podcast guesting, apple podcasts scraper.

### Why the email hit-rate is near 100%

Discovery uses the keyless **iTunes Search API** (returns the RSS `feedUrl` + metadata).
Apple **mandates** an owner email in every podcast RSS feed:

```xml
<itunes:owner><itunes:email>host@show.com</itunes:email><itunes:name>Host</itunes:name></itunes:owner>
```

so the host/business email is public and present for ~all shows. No login, no anti-bot,
**no residential proxy** — the iTunes API and RSS feeds are open. Validated live: 6/6 shows
returned a real host email on the first sample.

### Input

| field | meaning |
|---|---|
| `searchTerms` | topics/keywords to discover podcasts by |
| `itunesIds` / `feedUrls` | enrich specific shows directly |
| `country` | Apple storefront (US, GB, DE, …) |
| `genre` | keep only shows whose primary genre matches |
| `maxPerTerm` | shows per keyword (1–200) |
| `minEpisodes` | drop inactive shows below this episode count |
| `requireEmail` | only return (and charge for) shows with a host email |
| `concurrency` | feeds fetched in parallel (default 8) |

### Output (one item per podcast)

```json
{
  "name": "The Personal Finance Podcast", "author": "Andrew Giancola",
  "ownerEmail": "giancola.andrew@gmail.com", "ownerName": "Andrew Giancola",
  "website": "https://mastermoney.co", "language": "en-us",
  "genre": "Investing", "genres": ["Business"], "episodeCount": 546,
  "itunesId": 1496681088, "itunesUrl": "https://podcasts.apple.com/...",
  "artworkUrl": "https://...", "lastReleaseDate": "2026-06-25T...", "sourceTerm": "personal finance"
}
```

### Pricing (PAY\_PER\_EVENT)

| event | suggested price | when |
|---|---|---|
| `podcast-scraped` | $0.03 | every podcast returned (host email almost always included) |

Single event keeps it simple — owner email is present for ~all shows. Undercuts
ryanclinton ($0.05/podcast) while platform cost is tiny (no residential proxy), so margin
is near-pure. Use `requireEmail: true` to only pay for shows that expose an email.

### Run locally

```bash
mkdir -p storage/key_value_stores/default
cp .actor/input.example.json storage/key_value_stores/default/INPUT.json
APIFY_LOCAL_STORAGE_DIR="$PWD/storage" \
  uv run --with apify --with "httpx[http2]" python -m src.main
```

### Deploy

```bash
apify push
```

# Actor input Schema

## `searchTerms` (type: `array`):

Discover podcasts by topic, e.g. "personal finance", "true crime". One search per item.

## `itunesIds` (type: `array`):

Enrich specific shows by their Apple Podcasts collection IDs.

## `feedUrls` (type: `array`):

Enrich specific shows directly by RSS feed URL.

## `country` (type: `string`):

2-letter country code for the iTunes storefront (US, GB, DE, ...).

## `genre` (type: `string`):

Keep only shows whose primary genre contains this text, e.g. "Business".

## `maxPerTerm` (type: `integer`):

How many shows to take from each search keyword (1-200).

## `minEpisodes` (type: `integer`):

Drop shows below this episode count (0 = no filter). Filters out inactive shows.

## `requireEmail` (type: `boolean`):

Skip (and do not charge for) shows whose RSS feed exposes no owner email.

## `concurrency` (type: `integer`):

Number of feeds fetched in parallel (1-30).

## `proxyConfiguration` (type: `object`):

Not required — iTunes API and RSS feeds are open. Add a proxy only for very high volume.

## Actor input object example

```json
{
  "searchTerms": [
    "personal finance"
  ],
  "country": "US",
  "maxPerTerm": 25,
  "minEpisodes": 0,
  "requireEmail": false,
  "concurrency": 8
}
```

# Actor output Schema

## `podcasts` (type: `string`):

Dataset items — one per show: name, author, ownerName, ownerEmail, genre, episodeCount, website, itunesUrl, feedUrl.

## `run` (type: `string`):

Public run page with status, stats and the results table.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "personal finance"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ritaya/podcast-email-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchTerms": ["personal finance"] }

# Run the Actor and wait for it to finish
run = client.actor("ritaya/podcast-email-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "personal finance"
  ]
}' |
apify call ritaya/podcast-email-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=ritaya/podcast-email-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/35fnqVZoFUYP6EtE9/builds/ygoDPl5o4h5kCmdGn/openapi.json
