# Brazil Business Directory Scraper (`nickslam/brazilyello-scraper`) Actor

Scrapes business directory data from BrazilYello (brazilyello.com) including company information, reviews, products, and more

- **URL**: https://apify.com/nickslam/brazilyello-scraper.md
- **Developed by:** [Nick](https://apify.com/nickslam) (community)
- **Categories:** Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 base results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 📁 Brazil Business Directory Scraper

### Extract Brazil business data with BrazilYello

[BrazilYello](https://www.brazilyello.com) (brazilyello.com) lists local businesses across Brazil. Use this Actor as a practical Brazil business directory API alternative: configure filters in the Input tab and download JSON, CSV, or Excel.

BrazilYello is built for Portuguese-language category and keyword discovery across Brazil’s largest metros—pair city filters with local search terms for better precision.

Runs on the **Apify platform**, so you get monitoring, scheduling, API access, dataset exports, and proxy support out of the box.

#### Features

| Capability | Details |
|------------|---------|
| Categories | e.g. restaurants, consultants, estate-agents |
| Keywords | e.g. "advogado", "clínica", "logística" |
| Locations | São Paulo, Rio de Janeiro, Belo Horizonte, Curitiba, Campinas, Porto Alegre, Goiânia, Brasília |
| Enrichment toggles | Contact, extended info, registration, media, reviews |
| Delivery | JSON / CSV / Excel + Apify API |

### Getting started: scrape Brazil companies

1. Open this Actor on Apify and go to the **Input** tab.
2. Add **categories**, **search keywords**, and/or **locations** relevant to Brazil (see schema enums).
3. Set **Max Results** (for example 20) for a first test run.
4. Enable only the data add-ons you need — more fields can move you into a higher pay-per-event tier.
5. Click **Start**, then download the dataset as JSON, CSV, or Excel when the run succeeds.

### Why collect Brazilian company data from brazilyello.com?

- Scale lead gen across Brazil’s largest metro markets
- Segment extracts by state capitals for regional sales teams
- Combine Portuguese category/keyword filters with city selection

### Input

Open the **Input** tab on the Actor page for schemas, enums, and defaults.

#### Search & filters

| Option | Description |
|--------|-------------|
| **Categories** | Business type slugs (e.g. restaurants, consultants, estate-agents) |
| **Search Keywords** | Free-text search (e.g. "advogado", "clínica", "logística") |
| **Locations** | Limit to Brazilian cities/areas: São Paulo, Rio de Janeiro, Belo Horizonte, Curitiba, Campinas, Porto Alegre, Goiânia, Brasília |

#### Scope & limits

| Option | Description |
|--------|-------------|
| **Max Results** | Cap on companies to collect (start small while testing) |
| **Max Pages** | Cap on listing pages crawled |
| **Start Page** | Resume from a later listing page if a previous run stopped early |

#### Optional data add-ons

| Option | What it adds |
|--------|---------------|
| **Include Contact & Details** | Hours, year established, employees, contact person |
| **Include Extended Company Information** | Website, description, maps coordinates |
| **Include Registration & Legal** | Registration / tax identifiers when shown |
| **Include Metadata** | Listing type, verified flag, years with platform, tags |
| **Include Media & Content** | Photos, logo, products/services |
| **Include Review Summary** | Rating average and review count |
| **Include Individual Reviews** | Full review text (slower; extra pages) |

### Output example

Always-on core fields when available:

- **Company name**
- **Company URL** – profile on BrazilYello
- **Phone number**
- **Address**
- **Categories**
- **Location**

```json
{
  "companyName": "Atlas Logística SP",
  "companyUrl": "https://www.brazilyello.com/company/example-profile",
  "phoneNumber": "+55 11 4000-2000",
  "location": "São Paulo"
}
```

### 💰 Pricing (Pay per event)

| Tier | Price per 1,000 results | When it applies |
|------|--------------------------|-----------------|
| **Base** | $0.50 | Core fields only; no search keywords; no individual reviews |
| **Extended** | $1.00 | Any extended / contact / registration / media / review-summary / metadata options |
| **Premium** | $2.00 | Search keywords **or** Include Individual Reviews |

You can cap spend per run with Apify’s max total charge setting.

### Tips

- **Rate limiting** – keep enabled to reduce blocking risk.
- **Individual reviews** – enable only when you truly need review text.
- **Start page** – useful to continue a large crawl.
- São Paulo and Rio runs grow quickly — set a modest Max Results while exploring categories.

### Other directory scrapers

| Actor | Link |
|-------|------|
| UAE (Yello.ae) | [Open on Apify](https://apify.com/nickslam/yello-ae-scraper) |
| Indonesia (IndonesiaYP) | [Open on Apify](https://apify.com/nickslam/indonesiayp-scraper) |
| Singapore (Yelu.sg) | [Open on Apify](https://apify.com/nickslam/yelu-sg-scraper) |
| Poland (PolandYP) | [Open on Apify](https://apify.com/nickslam/poland-yp-scraper) |

### Is it legal to scrape BrazilYello?

This Actor collects information that businesses have published on BrazilYello. Results may include personal data (for example a contact person). You are responsible for using the data in line with applicable laws (including GDPR where relevant) and BrazilYello’s terms. Use the scraper only for legitimate purposes; if unsure, consult legal counsel. For feedback or bugs, use the Actor’s **Issues** tab on Apify.

# Actor input Schema

## `categories` (type: `array`):

Optional: Filter by business categories.

Example: For category page like `https://www.brazilyello.com/category/Estate_agents` enter only `Estate_agents`.

## `searchKeywords` (type: `string`):

Optional: Search for specific companies or keywords. Enables premium pricing.

## `locations` (type: `array`):

Optional: Filter by Brazil cities/regions. If not provided, all cities will be scraped.

## `maxResults` (type: `integer`):

Maximum number of companies to scrape

## `startPage` (type: `integer`):

Page number to start scraping from. Useful for resuming previous searches.

## `maxPages` (type: `integer`):

Maximum number of pages to scrape. Each page usually contains around 20 companies.

## `enableRateLimiting` (type: `boolean`):

Enable rate limiting (2 seconds delay) or disable (0.5 seconds delay)

## `hybridStrategy` (type: `string`):

`fast` = Cheerio + Impit HTTP (default). `browser` = Playwright (fallback).

## `includeContactDetails` (type: `boolean`):

Include business hours, year established, employees, contact person

## `includeExtendedCompanyInfo` (type: `boolean`):

Include website, description, and maps coordinates

## `includeRegistrationLegal` (type: `boolean`):

Include registration number and VAT number

## `includeMetadata` (type: `boolean`):

Include listing type, verified status, years with platform, tags

## `includeMediaContent` (type: `boolean`):

Include photos, logo, products

## `includeReviewSummary` (type: `boolean`):

Include review summary (rating and count) - Fast extraction

## `includeIndividualReviews` (type: `boolean`):

Include individual reviews - Slower (requires additional pages). Enables premium pricing.

## `proxy` (type: `object`):

Select proxies to be used by the crawler. Recommended: Use Apify Proxy with Residential proxies for better Cloudflare bypass.

## Actor input object example

```json
{
  "categories": [
    "restaurants",
    "consultants"
  ],
  "searchKeywords": "restaurant, IT services",
  "locations": [
    "Sao_Paulo",
    "Rio_de_Janeiro"
  ],
  "maxResults": 20,
  "startPage": 1,
  "maxPages": 5,
  "enableRateLimiting": true,
  "hybridStrategy": "fast",
  "includeContactDetails": false,
  "includeExtendedCompanyInfo": false,
  "includeRegistrationLegal": false,
  "includeMetadata": false,
  "includeMediaContent": false,
  "includeReviewSummary": false,
  "includeIndividualReviews": false,
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "BR"
  }
}
```

# Actor output Schema

## `companyData` (type: `string`):

Dataset containing all scraped company information

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "categories": [
        "restaurants"
    ],
    "locations": [
        "Sao_Paulo"
    ],
    "proxy": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "BR"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("nickslam/brazilyello-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "categories": ["restaurants"],
    "locations": ["Sao_Paulo"],
    "proxy": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "BR",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("nickslam/brazilyello-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "categories": [
    "restaurants"
  ],
  "locations": [
    "Sao_Paulo"
  ],
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "BR"
  }
}' |
apify call nickslam/brazilyello-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=nickslam/brazilyello-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HuNMC2SNdFZCdNpoc/builds/9AiM05UyQxHsD2Wwc/openapi.json
