# DACH Impressum Scraper — Emails, Phone, VAT & Directors (`scrapersdelight/imprint-contact-scraper`) Actor

Turn a list of DE/AT/CH company domains into firmographic B2B leads from each site's legally-mandated Impressum: email, phone, address, managing director, register court + HRB/HRA/FN, and VAT ID. Full DACH coverage. No login.

- **URL**: https://apify.com/scrapersdelight/imprint-contact-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Automation, Lead generation, Agents
- **Stats:** 8 total users, 4 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$15.00 / 1,000 per imprint lead returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## DACH Impressum Scraper — Emails, Phone, VAT & Directors

Turn a plain list of **German, Austrian and Swiss company domains** into complete, sellable B2B leads — scraped straight from each site's legally-mandated **Impressum** (legal notice). Paste your domains, get back a clean firmographic record per company.

### What it does

Every commercial DE/AT/CH website must publish an Impressum (§5 TMG/DDG in Germany, §5 ECG in Austria, plus the Swiss disclosure rules). That page is a goldmine of structured company data. For each domain you give it, this Actor:

1. Fetches the homepage and **discovers the Impressum / legal-notice link** (no fragile URL guessing).
2. Fetches that page and **extracts the firmographic fields** with DACH-tuned parsers.
3. Returns one flat lead row — ready for your CRM, outreach tool or enrichment pipeline.

### Why imprints are the cleanest EU lead source

Impressum pages are **legally-required public disclosures** — the company itself publishes its contact and registration details by law. That makes them the most **GDPR-defensible** starting point for EU B2B prospecting: no logins, no scraped private profiles, just statutory public data. And unlike a Maps listing, an imprint gives you the **managing director, commercial-register entry and VAT ID** — the fields that identify the actual legal entity.

### Full DACH coverage

Most imprint scrapers are Germany-only. This one handles **DE, AT and CH**: it rotates proxy IPs across the three countries, parses German-, Austrian- and Swiss-format phone numbers, and recognises `DE`, `ATU` and `CHE` VAT IDs plus `HRB` / `HRA` / `FN` / `CHE` register numbers.

### Input

| Field | What it is |
|---|---|
| **Company domains or imprint URLs** | One entry per company — bare domain (`trigema.de`), homepage URL, or a direct Impressum URL. One lead per domain. |
| **Country** | `Auto` (rotate & infer) or force `DE` / `AT` / `CH`. |
| **Max results** | Cap on leads this run. `0` = the whole list. |
| **Proxy** | Apify's shared proxy pool by default (rotates DACH countries, retries a fresh IP on any block). Switch to RESIDENTIAL in one click if a target domain sits behind an enterprise WAF. |

Example:

```json
{ "domains": ["trigema.de", "mymuesli.com", "example-gmbh.at"], "countryHint": "auto", "maxItems": 1000 }
```

### Output

One record per domain:

| Field | Description |
|---|---|
| `domain` | The input company domain. |
| `imprintUrl` | The discovered Impressum page. |
| `companyName` | Company / site name. |
| `email` | Contact email (prefers `mailto:`). |
| `phone` | Contact phone (prefers `tel:`). |
| `street`, `postalCode`, `city`, `country` | Postal address. |
| `managingDirector` | Geschäftsführer / Vorstand / Inhaber. |
| `registerCourt` | Amtsgericht / Registergericht / Firmenbuchgericht. |
| `registerNumber` | HRB / HRA / FN / CHE commercial-register number. |
| `vatId` | USt-IdNr / ATU / CHE VAT ID. |
| `website` | The company website. |
| `socialProfiles` | LinkedIn / Xing / Facebook / Instagram / X / YouTube links. |
| `scrapedAt` | ISO timestamp. |
| `status` | `ok` · `no_imprint_found` · `blocked` · `error`. |

### Pricing

**Pay per result.** You are charged a single event **only for each successfully parsed imprint** (`status: ok`). Domains that are blocked, have no discoverable imprint, or error out are still returned so you can see them — but they are **never charged**.

### Limits

- **Bring your own domains.** This is not a keyword-discovery tool — it does not find companies for you; it enriches the domain list you supply.
- **Enterprise-WAF sites may block.** A minority of large brands sit behind aggressive bot walls (Cloudflare captcha, 403). Those come back `status: blocked` and are not charged. SMB imprints — the target lead — are clean, plain HTML.
- **Extraction is best-effort.** Imprint layouts vary; a field the page doesn't publish comes back `null`. Email + at least one firmographic field (phone / VAT / register / director) is the norm for a clean SMB imprint.

### Legal

This Actor reads only **statutory public disclosures** (Impressum / legal-notice pages that companies are legally required to publish). You are responsible for complying with each site's Terms of Service and with all applicable outreach/marketing law (GDPR, UWG, etc.) when contacting the leads.

# Actor input Schema

## `domains` (type: `array`):

One entry per company. Paste bare domains (e.g. "trigema.de"), full homepage URLs ("https://www.mymuesli.com") or direct Impressum URLs. The Actor discovers each site's Impressum/legal-notice page and extracts the firmographic contact fields. One lead per domain.

## `countryHint` (type: `string`):

DACH country of the domains. Biases the proxy country and phone/VAT parsing. "Auto" rotates DE/AT/CH residential IPs and infers the country from the VAT ID / TLD.

## `maxItems` (type: `integer`):

Cap on domains processed / leads returned this run (cost & speed guard). Default 1000; set 0 for the whole input list.

## `requestConcurrency` (type: `integer`):

Max domains fetched in parallel. Higher = faster; keep modest to respect the sites.

## `proxyConfiguration` (type: `object`):

Proxy settings. The default is Apify's shared proxy pool, and the Actor retries a fresh IP on any soft-block. Measured on 2026-08-02 over 18 DACH domains x 5 runs on the shared pool vs 2 residential control runs: identical per-domain results on every field the leads are built from (email, phone, VAT ID, managing director, register court/number, address). RESIDENTIAL is a one-click escalation for enterprise-WAF domains.

## Actor input object example

```json
{
  "domains": [
    "trigema.de",
    "mymuesli.com"
  ],
  "countryHint": "auto",
  "maxItems": 1000,
  "requestConcurrency": 10,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `records` (type: `string`):

The dataset of scraped DACH company leads (one item per domain).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "trigema.de",
        "mymuesli.com"
    ],
    "countryHint": "auto",
    "maxItems": 1000,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/imprint-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "trigema.de",
        "mymuesli.com",
    ],
    "countryHint": "auto",
    "maxItems": 1000,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/imprint-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "trigema.de",
    "mymuesli.com"
  ],
  "countryHint": "auto",
  "maxItems": 1000,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call scrapersdelight/imprint-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scrapersdelight/imprint-contact-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/zXxZknArRHaRzJ8Ef/builds/GNhyZdDhVD7cDCo0K/openapi.json
