# YellowPages Scraper - Local Business Leads to CSV/Excel (`matrix-crawl/yellowpages-scraper`) Actor

Scrape YellowPages.com business listings by search term and location. Extract name, phone, address, website, rating, reviews, categories and years in business into CSV, Excel or JSON. Optional full-profile mode adds description, geo, hours and social links. No login, no API key.

- **URL**: https://apify.com/matrix-crawl/yellowpages-scraper.md
- **Developed by:** [Matrix Crawl](https://apify.com/matrix-crawl) (community)
- **Categories:** Lead generation, SEO tools, Automation
- **Stats:** 2 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## YellowPages Scraper — Local Business Leads (Phone, Address, Website) to CSV/Excel

**Fast YellowPages.com scraper for local-business lead generation.** Search any business type in any US city and pull clean, spreadsheet-ready rows — name, phone, address, website, rating, reviews, categories, years in business — with no login and no API key.

`yellowpages scraper` · `yellowpages.com scraper` · `local business leads` · `b2b lead generation` · `business directory` · `phone number scraper` · `local seo` · `sales prospecting`

Built for sales teams building local lead lists, agencies doing local SEO, and analysts mapping a market by city and category.

***

### 👥 Who Is This For?

- **B2B Sales / Lead-gen:** Build targeted call/email lists by trade + city.
- **Local SEO Agencies:** Pull local competitors with ratings, categories, and websites.
- **Market Analysts:** Map a category across geographies (counts, ratings, tenure).

***

### ⚡ Features

- **Contact data** — business name, phone, street address, locality, website.
- **Reputation** — star rating and review count.
- **Firmographics** — categories, years in business.
- **Multi-search** — several search terms in one run, one location filter.
- **Full-profile mode (optional)** — description, address parts, geo lat/lng, founding year, opening hours, social links.
- **De-duplicated** — by YellowPages profile URL.
- **Cost control** — hard `maxItems` cap.
- **Export** — CSV/Excel/JSON, or call via Apify's MCP server.

***

### 🚀 How to Use

1. Add one or more **search terms** (e.g. `plumber`, `dentist`, `roofing contractor`).
2. Set a **location** (e.g. `New York, NY` or a ZIP like `90210`).
3. Set **Maximum Items** and press **Start**.
4. Export as CSV, Excel, JSON, or HTML.

#### Sample Input

```json
{
  "searchTerms": ["plumber", "electrician"],
  "location": "New York, NY",
  "maxPagesPerSearch": 3,
  "maxItems": 500,
  "scrapeProfileDetails": false,
  "proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}
```

***

### 🔎 Two modes

- **Directory (default):** everything from the results page — name, phone, address, website, rating, reviewCount, categories, yearsInBusiness, snippet. One request per results page.
- **Full profile (`scrapeProfileDetails: true`):** visits each business profile for the rich object — full `description`, `address` object (street/city/region/zip/country), `geo` (lat/lng), `foundingDate`, `openingHours`, `socialLinks`, `areaServed`. **+1 request per business** (slower, more bandwidth). Businesses whose profile is temporarily blocked still return their directory row — no data is lost.

***

### 📊 Output Data Dictionary (Directory mode)

| Field | Type | Example |
|-------|------|---------|
| `name` | string | `"RR Plumbing Roto-Rooter"` |
| `phone` | string | null | `"(917) 997-7507"` |
| `streetAddress` | string | null | `"450 7th Ave"` |
| `locality` | string | null | `"New York, NY 10001"` |
| `website` | string | null | `"https://www.rotorooter.com/manhattan/"` |
| `rating` | number | null | `4.5` |
| `reviewCount` | integer | null | `28` |
| `categories` | string | null | `"Plumbers, Plumbing-Drain & Sewer Cleaning"` |
| `yearsInBusiness` | integer | null | `91` |
| `snippet` | string | null | `"From Business: When you need an emergency plumber…"` |
| `yellowPagesUrl` | string | null | `"https://www.yellowpages.com/…/mip/…"` |
| `searchTerm` | string | `"plumber"` |
| `location` | string | `"New York, NY"` |
| `scrapedAt` | string (ISO) | `"2026-07-01T10:00:00.000Z"` |

> Fields not shown on a listing come back `null`.

#### Extra fields with **Full profile** mode (`scrapeProfileDetails: true`)

| Field | Type | Description |
|-------|------|-------------|
| `description` | string | Full business description |
| `address` | object | street, locality, region, postalCode, country |
| `geo` | object | `{ lat, lng }` |
| `foundingDate` | string | Founding year, e.g. `"1935"` |
| `openingHours` | array | string | Opening hours |
| `areaServed` | string | Service area, e.g. `"New York, NY"` |
| `socialLinks` | array | Profile links (LinkedIn, Twitter, YouTube…) |
| `email` | string | null | Email if listed (often absent) |
| `categoriesFull` | string | Full category list from the profile |

***

### 💡 Proxy (required)

YellowPages blocks data-center IPs, so this actor uses **residential** proxy.

- **Default:** Apify residential — works out of the box.
- **Optional:** plug in your own residential proxy for higher volume.

Tip: use `maxItems` / `maxPagesPerSearch` to control cost.

### ⚖️ Compliance & Responsible Use

Collects **publicly listed business** directory data only. You are responsible for lawful use, anti-spam compliance for any outreach, and respecting YellowPages.com's Terms of Service and `robots.txt`.

# Actor input Schema

## `searchTerms` (type: `array`):

Business types / keywords to search on YellowPages, e.g. "plumber", "dentist", "roofing contractor".

## `location` (type: `string`):

City/state or ZIP to search in, e.g. "New York, NY" or "90210".

## `maxPagesPerSearch` (type: `integer`):

How many result pages to crawl per search term (30 listings per page).

## `maxItems` (type: `integer`):

Hard cap on total businesses scraped (cost control). 0 = no cap.

## `scrapeProfileDetails` (type: `boolean`):

Visit each business profile page for extra fields (full description, address parts, geo lat/lng, founding year, opening hours, social links, area served). Slower and uses more proxy bandwidth — one extra request per business.

## `debugLog` (type: `boolean`):

If a page returns 200 but parses 0 rows, save the HTML to the key-value store for inspection.

## `proxyConfiguration` (type: `object`):

Residential proxy is required — YellowPages blocks data-center IPs. Default Apify residential works out of the box.

## Actor input object example

```json
{
  "searchTerms": [
    "plumber"
  ],
  "location": "New York, NY",
  "maxPagesPerSearch": 1,
  "maxItems": 20,
  "scrapeProfileDetails": false,
  "debugLog": false,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "plumber"
    ],
    "location": "New York, NY"
};

// Run the Actor and wait for it to finish
const run = await client.actor("matrix-crawl/yellowpages-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["plumber"],
    "location": "New York, NY",
}

# Run the Actor and wait for it to finish
run = client.actor("matrix-crawl/yellowpages-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "plumber"
  ],
  "location": "New York, NY"
}' |
apify call matrix-crawl/yellowpages-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=matrix-crawl/yellowpages-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/oLXOvgpphhOuxfgIN/builds/R1P3HI7XcUVINfLai/openapi.json
