# Website Contact Scraper — Emails, Phones & Socials (`hipersoft/website-contact-scraper`) Actor

Extract business emails, phone numbers and social media profiles from any list of websites. De-obfuscates emails, reads Organization schema, and optionally crawls contact/about pages. Bulk-ready, no login or API key.

- **URL**: https://apify.com/hipersoft/website-contact-scraper.md
- **Developed by:** [hiper soft](https://apify.com/hipersoft) (community)
- **Categories:** Lead generation, Automation
- **Stats:** 2 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.0024 / website scraped

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Website Contact Scraper — Emails, Phones & Social Profiles

Turn a list of websites into a clean contact sheet. Give the Actor any set of
domains or URLs and it returns **emails, phone numbers, and social media profiles**
for each one — ready for lead lists, CRM enrichment, or outreach.

No login, no API key. Bulk-ready: feed it thousands of sites at once.

### What it extracts

| Field | Notes |
|---|---|
| `emails` | From `mailto:` links, page text, and **de-obfuscated** forms (`name [at] site [dot] com`). Junk/placeholder addresses filtered out. |
| `phones` | Real declared numbers only — from `tel:` links and `Organization` schema — normalized (kept in E.164-ish form). We deliberately skip loose page digits so you never get dates/IDs masquerading as phones. |
| `socials` | Classified links: LinkedIn, Twitter/X, Facebook, Instagram, YouTube, TikTok, Pinterest, GitHub, Telegram, WhatsApp. Share/intent URLs excluded. |
| `companyName` | From the site's `Organization`/`LocalBusiness` JSON-LD when present. |
| `address` | Postal address from structured data when present. |
| `domain`, `url` | Canonical domain and the final URL after redirects. |
| `pagesCrawled` | How many pages were read for this site. |

#### Why it's more complete

Unlike a homepage-only scraper, it optionally **crawls the contact/about/imprint
pages** on the same site, reads **JSON-LD structured data** (which many sites use to
publish their real email, phone, address and social links), and **de-obfuscates**
emails that are deliberately hidden from naive scrapers.

### Input

```json
{
  "websites": ["apify.com", "stripe.com", "https://example.org"],
  "deepCrawl": true,
  "maxPagesPerSite": 5,
  "maxConcurrency": 10
}
```

- **websites** — domains or URLs, one per line. (Or use **startUrls** for a request list.)
- **deepCrawl** — follow contact/about pages for more results (default on). Turn off for one request per site.
- **maxPagesPerSite** — cap pages fetched per site (default 5).
- **maxConcurrency** — sites processed in parallel (default 10).
- **maxItems** — cap total sites (0 = all).
- **proxyConfiguration** — optional; enable if you hit blocks.

### Output (one record per website)

```json
{
  "input": "apify.com",
  "url": "https://apify.com/",
  "domain": "apify.com",
  "companyName": "Apify",
  "emails": ["hello@apify.com", "support@apify.com"],
  "phones": ["+14155551234"],
  "socials": {
    "linkedin": "https://www.linkedin.com/company/apifytech",
    "twitter": "https://twitter.com/apify",
    "github": "https://github.com/apify"
  },
  "address": "...",
  "pagesCrawled": 3,
  "error": null
}
```

Sites where nothing was found still return a row, with `error` describing why
(e.g. a network block) or `null` when the site simply had no public contacts.

### Pricing

Pay per website result. No monthly fee — you only pay for what you run.

### FAQ

**Do I need an API key or login?**
No. There's no login and no API key — paste your list of domains or URLs and run. It's bulk-ready, so you can feed it thousands of sites at once.

**How many leads can I get per run?**
One record per website, with no fixed ceiling — cap the total with `maxItems` (0 = all) and control depth per site with `maxPagesPerSite`. Feed thousands of domains in a single run.

**Is scraping website contact info legal?**
The Actor extracts only publicly available contact details already published on the target sites. Collecting public data is generally permissible, but you are responsible for using it in line with GDPR, CAN-SPAM and each site's terms.

**Does it collect emails and contact info?**
Yes. It pulls emails (including de-obfuscated `name [at] site [dot] com` forms), declared phone numbers, social profiles, company name and postal address from page text and JSON-LD structured data.

**What's the output format?**
One structured JSON record per website with `emails`, `phones`, `socials`, `companyName`, `address` and crawl metadata. Sites with no contacts still return a row. Export as JSON, CSV or Excel.

### Related Actors

Combine website enrichment with these lead-generation companions:

- [Google Maps Email Extractor](https://apify.com/hipersoft/google-maps-email-extractor) — find local businesses on Google Maps and extract their emails and socials in one pass.
- [Google Maps Scraper](https://apify.com/hipersoft/google-maps-scraper) — get business listings with websites to feed into this Actor.
- [Clutch.co Agency Scraper](https://apify.com/hipersoft/clutch-scraper) — pull B2B agency profiles with ratings and budget signals.
- [Website Content Crawler](https://apify.com/hipersoft/website-content-crawler) — extract clean page text and metadata from the same sites.

### Notes & fair use

Extracts only **publicly available** contact information already published on the
target websites. You are responsible for using the data in line with applicable
laws (GDPR/CAN-SPAM etc.) and each site's terms.

# Actor input Schema

## `websites` (type: `array`):

List of website URLs or domains to extract contacts from (e.g. "stripe.com", "https://apify.com"). One per line.

## `startUrls` (type: `array`):

Alternative to "Websites": a request list of {"url": "..."} objects, e.g. from a linked dataset or Key-Value store.

## `deepCrawl` (type: `boolean`):

When on, follow up to a few likely contact/about/imprint pages on the same site for more complete results. Turn off for one-request-per-site (cheaper, faster).

## `maxPagesPerSite` (type: `integer`):

Upper bound on pages fetched per website (homepage + contact pages).

## `maxConcurrency` (type: `integer`):

How many websites to process in parallel.

## `maxItems` (type: `integer`):

Cap on how many input websites to process (0 = no cap, process all).

## `proxyConfiguration` (type: `object`):

Optional. Most sites work from datacenter IPs; enable a proxy if you hit blocks or need a specific country.

## Actor input object example

```json
{
  "websites": [
    "apify.com"
  ],
  "deepCrawl": true,
  "maxPagesPerSite": 5,
  "maxConcurrency": 10,
  "maxItems": 0,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "apify.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("hipersoft/website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": ["apify.com"] }

# Run the Actor and wait for it to finish
run = client.actor("hipersoft/website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "apify.com"
  ]
}' |
apify call hipersoft/website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=hipersoft/website-contact-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/COX3ExTTWZ8jQwJ2Y/builds/DSnd3VdsH4sk66WYO/openapi.json
