# Yellow Pages Scraper USA — Business Leads (`muhammadafzal/yellow-pages-us-scraper`) Actor

Scrape Yellow Pages USA business listings with names, phones, websites, addresses, ratings, reviews, and structured local lead data.

- **URL**: https://apify.com/muhammadafzal/yellow-pages-us-scraper.md
- **Developed by:** [Muhammad Afzal](https://apify.com/muhammadafzal) (community)
- **Categories:** Lead generation, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 business listing extracteds

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Yellow Pages Scraper USA — Business Leads

Extract US business listings from **yellowpages.com** at scale. Get business name, phone, address, website, email, hours, services, categories, ratings, review counts, social links, BBB accreditation, claimed status, years in business, and optional customer reviews — all in clean, CRM-ready structured JSON.

Search by keyword + location, paste category URLs, or pull individual business profiles. Mix all three input modes in a single run. Optimized for B2B lead generation, local-SEO analysis, market research, and AI-agent pipelines.

### What it does

- **Search mode**: provide `searchTerms` + `locations` (cross-joined) — e.g. `["plumber", "electrician"]` × `["Boston, MA", "Austin, TX"]` = 4 searches
- **URL mode**: paste yellowpages.com search URLs, city-category pages (`/boston-ma/plumbers`), or individual business profiles (`/mip/<slug>-<id>`)
- **Multi-input**: combine search terms and direct URLs in one run
- **Pagination**: crawls up to `maxPagesPerSearch` result pages per search
- **Detail enrichment**: visits each business profile page for the full field set (email, hours, services, socials, trust signals)
- **Reviews**: optional per-business customer reviews (newest-first, capped via `maxReviewsPerBusiness`)
- **Sort**: by relevance (default), distance, name, or rating

### Use cases

- Build targeted B2B prospect lists by category and ZIP/city for outbound sales
- Pull every business in a category or metro for market coverage and competitor mapping
- Filter for unverified/owner-unclaimed listings to identify outreach opportunities
- Enrich existing CRM data with verified phone, website, structured address, hours, socials
- Extract customer reviews for sentiment analysis, quality benchmarking, reputation research
- Map local business density and compare market-level metrics (ratings, review volume, years in business)
- Seed RAG/LLM corpora with structured local-business records

### Input

| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| `searchTerms` | string\[] | no | `["plumber"]` | Keywords/categories — cross-joined with `locations` |
| `locations` | string\[] | no | `["Boston, MA"]` | US locations in "City, ST" or ZIP format |
| `startUrls` | array | no | `[]` | Direct yellowpages.com URLs (search, category, or profile) |
| `sortBy` | enum | no | `default` | `default`/`distance`/`name`/`rating` |
| `includeReviews` | boolean | no | `false` | Embed customer reviews per business |
| `maxReviewsPerBusiness` | integer | no | `30` | Cap on reviews per business (0 = all, hard cap ~2000) |
| `maxResults` | integer | no | `100` | Hard cap on total business rows (0 = unlimited, internal cap 100k) |
| `maxPagesPerSearch` | integer | no | `5` | Max result pages per search/URL |
| `proxyConfiguration` | object | no | US residential | **Required**: US residential proxies (Cloudflare blocks datacenter) |

### Output

One record per business in the dataset:

| Field | Type | Example |
|---|---|---|
| `name` | string | "Acme Plumbing" |
| `phone` | string|null | "(617) 555-1234" |
| `email` | string|null | "info@acme.com" |
| `website` | string|null | "https://acme.com" |
| `address` | object | `{street, city, state, postalCode, country}` |
| `addressFormatted` | string | "123 Main St, Boston, MA 02108" |
| `coordinates` | object|null | `{lat, lng}` |
| `categories` | string\[] | `["Plumbing","Contractor"]` |
| `primaryCategory` | string|null | "Plumbing" |
| `services` | string\[] | `["Drain cleaning","Water heater repair"]` |
| `paymentMethods` | string\[] | `["Visa","Mastercard"]` |
| `hours` | object|null | `{mon:"8:00-17:00",...}` |
| `rating` | number|null | 4.5 |
| `reviewCount` | number|null | 128 |
| `reviews` | array|null | `[{author,rating,date,text}]` (when `includeReviews=true`) |
| `socialLinks` | object | `{facebook,instagram,linkedin,twitter,youtube}` |
| `logo` | string|null | url |
| `photos` | string\[] | `[url,...]` |
| `claimed` | boolean | true |
| `bbbAccredited` | boolean | false |
| `yearsInBusiness` | number|null | 12 |
| `description` | string|null | "Family-owned plumbing services since..." |
| `tagline` | string|null | "Your trusted local plumber" |
| `isAd` | boolean | false |
| `yellowPagesId` | string | "123456789" |
| `profileUrl` | string | `https://www.yellowpages.com/.../mip/acme-12345` |
| `sourceUrl` | string | URL crawled |
| `searchTerm` | string | "plumber" |
| `searchLocation` | string | "Boston, MA" |
| `scrapedAt` | string | ISO 8601 timestamp |

### Pricing

This actor uses **pay-per-event** pricing — you only pay for results delivered.

| Event | Price | When charged |
|---|---|---|
| Actor Start | $0.00005 | Once per run (per 1 GB memory) |
| Business Listing | $0.001 | Per business record extracted |

**Cost examples:**

- 100 businesses = ~$0.10
- 1,000 businesses = ~$1.00
- 10,000 businesses = ~$10.00

Usage-based billing (compute + proxy passthrough) is also enabled for heavy users — see the Pricing tab for tier details.

### Proxy requirements

**US residential proxies are required.** Yellow Pages is fronted by Cloudflare and blocks datacenter IPs and non-US residential ranges. The default proxy configuration uses Apify's `RESIDENTIAL` group with `countryCode: US`.

For aggressive anti-bot environments, you can supply a custom premium residential proxy (Decodo, Bright Data, IPRoyal) via the proxy configuration input.

### Quick start

```bash
## Install Apify CLI
npm install -g apify-cli

## Clone and deploy
cd yellow-pages-us-scraper
npm install
apify login
apify push -f -w 120
```

### API usage

```bash
## Run via API
curl -X POST "https://api.apify.com/v2/acts/USERNAME~yellow-pages-us-scraper/runs?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"searchTerms":["dentist"],"locations":["Austin, TX"],"maxResults":50}'
```

```javascript
// Node.js / Apify SDK
import { Actor } from 'apify';

const run = await Actor.call('USERNAME/yellow-pages-us-scraper', {
    searchTerms: ['dentist'],
    locations: ['Austin, TX'],
    maxResults: 50,
    includeReviews: true,
    maxReviewsPerBusiness: 10,
});

const { items } = await Actor.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

### Tips

- **Start small**: set `maxResults: 10` to verify data quality before scaling
- **Reviews are slow**: enabling `includeReviews` roughly doubles runtime and cost — use only when needed
- **Large coverage**: split a big request across multiple cities/ZIPs — a single search caps at ~6,000 results
- **Sort by rating**: set `sortBy: "rating"` to surface top-rated businesses first
- **Cross-join math**: 2 terms × 3 locations = 6 searches; results are merged and deduplicated by profile URL

### Limitations

- Yellow Pages itself caps results at roughly 30 pages (~6,000 listings) per search term + location pair
- Reviews are capped at ~2,000 per business by Yellow Pages
- Email is only available when present on the business profile page (many YP listings do not show email)

### Export & integrations

Export scraped data to JSON, CSV, or Excel from the Apify Console. Use webhooks or the Apify API to push results to HubSpot, Salesforce, Google Sheets, Airtable, or any HTTP endpoint. Schedule runs via the Apify Scheduler for ongoing monitoring.

***

Export scraped data, run the scraper via API, schedule and monitor runs, or integrate with other tools.

### What is Yellow Pages Scraper USA?

**Yellow Pages Scraper USA** turns the target data into structured, reusable results on Apify. Use it when you need repeatable collection for sales teams, agencies, recruiters, market researchers, and data-enrichment workflows without maintaining a custom scraper or one-off integration. Run it manually, schedule recurring jobs, call it through the Apify API, or connect it to an AI agent through the Apify MCP server.

The Actor stores results in an Apify dataset, where they can be previewed and exported as JSON, CSV, Excel, XML, or RSS. Availability and completeness depend on the source, supplied inputs, public visibility, authentication requirements, and upstream rate limits.

### Use cases for Yellow Pages Scraper USA

- Build structured datasets for research, reporting, enrichment, or monitoring.
- Automate repetitive collection with schedules, webhooks, and API calls.
- Feed clean records into spreadsheets, databases, CRMs, BI tools, AI agents, or RAG pipelines.
- Track changes over time by running the same validated input on a schedule.
- Replace fragile manual copy-and-paste work with a reproducible Apify workflow.

### How to use Yellow Pages Scraper USA

1. Open the Actor input page and choose a focused, valid target.
2. Set a conservative result limit for the first run.
3. Start the Actor and inspect the dataset for coverage and field availability.
4. Export the results or connect the dataset to your downstream system.
5. Scale gradually and use scheduling, pagination, or proxies when supported.

#### Important input options

- `searchTerms` — Keywords or business categories to search for on Yellow Pages. Examples: "plumber", "dentist", "italian restaurant", "auto repair". Each term is cross-joined with every entry in Locations —
- `locations` — US locations in "City, ST" format, or 5-digit ZIP codes. Examples: "Boston, MA", "90210", "Brooklyn, NY", "Austin, TX". Each location is cross-joined with every search term. Use "City, ST" f
- `startUrls` — Paste one or more yellowpages.com URLs. Accepts search results (yellowpages.com/search?...), city-category pages (yellowpages.com/boston-ma/plumbers), or individual business profiles (yellow
- `sortBy` — How Yellow Pages sorts the search results before we collect them. "default" = relevance (YP's own ranking). "distance" = nearest first. "name" = alphabetical. "rating" = highest rated first.
- `includeReviews` — When enabled, embed each business's customer reviews on the result row (newest first). Off by default — turning this on makes runs slower and more expensive, since each business needs additi
- `maxReviewsPerBusiness` — Cap on the number of reviews captured per business when "Include Client Reviews" is on (most recent first). Set to 0 to capture every available review (an internal hard cap of ~2,000 reviews
- `maxResults` — Hard cap on total business rows across all searches and URLs. Default 100 — increase for bigger runs, or set to 0 for no cap (an internal upper limit of 100,000 still applies). The actor sto
- `maxPagesPerSearch` — Maximum result pages to crawl per (search term x location) pair or per start URL. Yellow Pages typically caps results around 30 pages. Lower this for faster, cheaper runs; raise it for exhau

### API and automation example

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('muhammadafzal/yellow-pages-us-scraper').call({
  // Add the same input fields you use in the Apify Console.
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

### Related Apify Actors

Use these dedicated tools when a neighboring data source or workflow is a better match:

- [Leads Finder Pro - B2B Leads with Emails \[Apollo Alternative\]](https://apify.com/muhammadafzal/leads-finder-pro)
- [Yellow Pages Australia Scraper — Business Leads & Reviews](https://apify.com/muhammadafzal/yellow-pages-au-scraper)
- [California CSLB Contractor License Scraper](https://apify.com/muhammadafzal/cslb-california-scraper)
- [Instagram Followers & Following Scraper — With Cookies](https://apify.com/muhammadafzal/instagram-following-scraper)
- [OpenTable Restaurants, Ratings & Reviews Scraper](https://apify.com/muhammadafzal/opentable-scraper)
- [Etsy Scraper Pro — Products, Prices, Reviews & Shop Data](https://apify.com/muhammadafzal/etsy-scraper-pro)
- [YouTube Thumbnail Downloader](https://apify.com/muhammadafzal/youtube-thumbnail-downloader)
- [AllTrails Scraper — Hiking Trails, Reviews & GPS Data](https://apify.com/muhammadafzal/alltrails-scraper)
- [Email Validator — Syntax, MX & Risk Checks](https://apify.com/muhammadafzal/email-address-validator)
- [Google Play Reviews Scraper](https://apify.com/muhammadafzal/google-play-reviews-scraper)

### Frequently asked questions

#### How many results can I scrape with Yellow Pages Scraper USA?

The practical total depends on the source, input limits, pagination, available records, run timeout, and upstream restrictions. Start with a small run, verify the output, and increase the limit gradually.

#### Can I integrate Yellow Pages Scraper USA with other apps?

Yes. Use Apify integrations, webhooks, schedules, dataset exports, Make, Zapier, Google Sheets, cloud storage, or your own application.

#### Can I use Yellow Pages Scraper USA with the Apify API?

Yes. Start runs with the Apify REST API or an official Apify client, then retrieve records from the run's default dataset. Keep your API token in a secret or environment variable.

#### Can I use Yellow Pages Scraper USA through an MCP Server?

Yes. The Apify MCP server can expose the Actor to compatible AI clients and agents. Review the input and expected cost before allowing an autonomous workflow to run it at scale.

#### Do I need proxies?

It depends on the source and volume. Use the default configuration first. For larger or geographically sensitive jobs, select an appropriate proxy configuration only when the Actor supports it.

#### Is it legal to scrape this data?

Scraping rules vary by source, jurisdiction, data type, and intended use. Collect only data you are authorized to access, respect applicable terms and privacy laws, and avoid restricted or personal data misuse. This documentation is not legal advice.

#### Your feedback

If a field is missing, a source layout has changed, or you need a supported use case documented, open an issue on the Actor page with a reproducible input and run ID.

# Actor input Schema

## `searchTerms` (type: `array`):

Keywords or business categories to search for on Yellow Pages. Examples: "plumber", "dentist", "italian restaurant", "auto repair". Each term is cross-joined with every entry in Locations — for example, 2 terms x 3 locations = 6 searches. Use this when the user describes a type of business or category. Use startUrls instead when the user provides specific yellowpages.com URLs.

## `locations` (type: `array`):

US locations in "City, ST" format, or 5-digit ZIP codes. Examples: "Boston, MA", "90210", "Brooklyn, NY", "Austin, TX". Each location is cross-joined with every search term. Use "City, ST" for best results — Yellow Pages geocodes the location and returns nearby businesses.

## `startUrls` (type: `array`):

Paste one or more yellowpages.com URLs. Accepts search results (yellowpages.com/search?...), city-category pages (yellowpages.com/boston-ma/plumbers), or individual business profiles (yellowpages.com/.../mip/<slug>-<id>). Filters in the URL are honored as-is. When provided alongside searchTerms, both are crawled. Use this when the user provides specific URLs instead of keywords.

## `sortBy` (type: `string`):

How Yellow Pages sorts the search results before we collect them. "default" = relevance (YP's own ranking). "distance" = nearest first. "name" = alphabetical. "rating" = highest rated first. Applies to keyword + location searches and category URLs; ignored for individual business profile URLs.

## `includeReviews` (type: `boolean`):

When enabled, embed each business's customer reviews on the result row (newest first). Off by default — turning this on makes runs slower and more expensive, since each business needs additional review-page fetches. Use maxReviewsPerBusiness to cap the count.

## `maxReviewsPerBusiness` (type: `integer`):

Cap on the number of reviews captured per business when "Include Client Reviews" is on (most recent first). Set to 0 to capture every available review (an internal hard cap of ~2,000 reviews per business still applies). Ignored when reviews are disabled.

## `maxResults` (type: `integer`):

Hard cap on total business rows across all searches and URLs. Default 100 — increase for bigger runs, or set to 0 for no cap (an internal upper limit of 100,000 still applies). The actor stops requesting new listing pages once this number is reached but keeps the full final page even if it slightly overshoots. A single search term + location pair caps at roughly 6,000 results (Yellow Pages itself rarely returns more) — split a large request across multiple cities or ZIPs to gather more.

## `maxPagesPerSearch` (type: `integer`):

Maximum result pages to crawl per (search term x location) pair or per start URL. Yellow Pages typically caps results around 30 pages. Lower this for faster, cheaper runs; raise it for exhaustive coverage. Default 5 balances coverage and cost.

## `proxyConfiguration` (type: `object`):

Apify proxy settings. US residential proxies (RESIDENTIAL, country US) are REQUIRED — Yellow Pages is fronted by Cloudflare and blocks datacenter IPs and non-US residential ranges. Leave as default unless you have a custom residential proxy provider (Decodo, Bright Data, IPRoyal).

## Actor input object example

```json
{
  "searchTerms": [
    "plumber"
  ],
  "locations": [
    "Boston, MA"
  ],
  "startUrls": [],
  "sortBy": "default",
  "includeReviews": false,
  "maxReviewsPerBusiness": 30,
  "maxResults": 10,
  "maxPagesPerSearch": 3,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Link to the dataset containing all extracted Yellow Pages business records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "plumber"
    ],
    "locations": [
        "Boston, MA"
    ],
    "startUrls": [],
    "sortBy": "default",
    "includeReviews": false,
    "maxReviewsPerBusiness": 30,
    "maxResults": 10,
    "maxPagesPerSearch": 3,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "US"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("muhammadafzal/yellow-pages-us-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["plumber"],
    "locations": ["Boston, MA"],
    "startUrls": [],
    "sortBy": "default",
    "includeReviews": False,
    "maxReviewsPerBusiness": 30,
    "maxResults": 10,
    "maxPagesPerSearch": 3,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "US",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("muhammadafzal/yellow-pages-us-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "plumber"
  ],
  "locations": [
    "Boston, MA"
  ],
  "startUrls": [],
  "sortBy": "default",
  "includeReviews": false,
  "maxReviewsPerBusiness": 30,
  "maxResults": 10,
  "maxPagesPerSearch": 3,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}' |
apify call muhammadafzal/yellow-pages-us-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=muhammadafzal/yellow-pages-us-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/dRXZfIbjxW9g6DhRx/builds/o0cNVCSPq8NptCHDd/openapi.json
