# Google Ads Transparency Scraper — Competitor Ads & Creatives (`silentshadow55/google-ads-transparency-actor`) Actor

Scrape every ad creative an advertiser runs on Google from the Ads Transparency Center. Resolve by advertiser name, ID, or domain; get image and HTML5 creatives with first/last-shown dates, region filtering, and full pagination — no 100-item cap. Cookieless, no login, no CAPTCHA.

- **URL**: https://apify.com/silentshadow55/google-ads-transparency-actor.md
- **Developed by:** [Esteban Ortega](https://apify.com/silentshadow55) (community)
- **Categories:** SEO tools, E-commerce
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 ad creatives

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Google Ads Transparency Scraper — Competitor Ads & Creatives

Scrape every ad creative an advertiser runs on Google straight from the **Google Ads Transparency Center** (adstransparency.google.com) — the fastest way to **download all of a competitor's Google ads**. Give it an advertiser name, an advertiser ID, or a domain, and get back a clean, structured list of that advertiser's image and HTML5 ad creatives — with first-shown / last-shown dates, region filtering, and **full pagination (no hidden 100-item cap)**. Cookieless, no login, no CAPTCHA.

### Why this one — complete captures or a loud error

| | This actor | Typical Transparency scraper |
|---|---|---|
| Pagination | Full token pagination — **every** creative | Often caps at ~the first page (~100) |
| When Google's payload shifts | **Loud, explicit error record** | Silent empty dataset |
| Per-region first/last-shown dates | Yes (opt-in enrichment) | Rare |
| Expiring asset URLs | Optional **rehost** persists the bytes | Links die |
| Partial runs | Marked `partial: true` + an error explaining where | Reported as complete |

The most-used Transparency scraper on the Store sits at ~2.2★ for exactly these reasons. This actor's contract is simple: **you get everything, or you get told loudly why not** — never a silent zero.

Built for competitive-ad researchers, performance marketers, brand-safety teams, and AI agents that need an advertiser's live ad library without the Google Ads Transparency Center's clunky UI.

### What it does

- **Resolves an advertiser three ways.** Look up by **advertiser name** ("Verizon"), by exact **advertiser ID** (`AR11385763688137883649`), or by **domain** ("verizon.com"). No fuzzy guesswork — the mode is explicit, and ambiguous names return disambiguation candidates you can pin.
- **Lists every creative with full token pagination.** Most scrapers silently stop at 100 results. This one follows Google's pagination token to the end (or your chosen cap), so you get the whole ad library — deduped by creative ID and guarded against stuck pagination tokens and runaway page loops.
- **Region filtering that actually works.** Pass a 2-letter country code (US, GB, CA…; UK/USA are aliased) to see only ads shown in that country. Uses Google's real geo IDs, validated against the live site. An unknown code fails fast with one error record instead of silently scraping the whole world.
- **Optional per-region enrichment.** Turn on creative enrichment to add every render variant and per-country last-shown dates for each creative.
- **Optional asset rehosting.** Google's `simgad` image and preview URLs expire fast — enable `downloadAssets` to persist the actual creative bytes in your run's key-value store.
- **Loud on anomalies, never silently empty.** If Google changes its internal payload shape or soft-blocks a request (the #1 reason competing scrapers return nothing), this actor surfaces an explicit error or notice record — including mid-pagination truncation (`partial: true` on the advertiser summary) — instead of pretending the result is a complete zero. A genuinely dormant advertiser (Google's own counts say 0 recent and 0 lifetime ads) is reported as a clean zero, not a false alarm.
- **Streaming results.** Every page of creatives is pushed to the dataset the moment it is scraped, so a timeout or migration late in a long run still leaves everything already fetched in your dataset.

### Output fields

Each creative is one dataset item (`type: "creative"`). Every run also emits one `advertiser` summary record per advertiser, plus `advertiser_candidate` / `domain_candidate` records for disambiguation. At most **one** `error` or `notice` record is emitted per advertiser target (and failures never consume your `maxCreativesPerAdvertiser` budget). All records share the identical key set, so exports stay rectangular.

Sample creative output item:

```json
{
  "type": "creative",
  "scraped_at": "2026-07-23T05:10:00+00:00",
  "advertiser_id": "AR11385763688137883649",
  "advertiser_name": "Verizon Value, Inc",
  "creative_id": "CR15308872423592427521",
  "format": "image",
  "format_enum": 1,
  "image_url": "https://tpc.googlesyndication.com/archive/simgad/6369736585639432871",
  "preview_url": null,
  "first_shown": "2025-09-22T07:00:00+00:00",
  "last_shown": "2026-07-23T04:14:25+00:00",
  "region": "US",
  "regions": [],
  "variants": [],
  "deep_link": "https://adstransparency.google.com/advertiser/AR11385763688137883649/creative/CR15308872423592427521?region=us",
  "asset_kvs_key": null,
  "error": null
}
```

Sample advertiser summary item:

```json
{
  "type": "advertiser",
  "advertiser_id": "AR11385763688137883649",
  "advertiser_name": "Verizon Value, Inc",
  "legal_name": "Verizon Value, Inc",
  "country": "US",
  "verified": true,
  "total_creatives": 500,
  "recent_ads": 500,
  "lifetime_ads": 600,
  "partial": false,
  "region": "US"
}
```

- `total_creatives` — the number of creatives **actually scraped in this run** (after your cap / region filter / any truncation).
- `recent_ads` / `lifetime_ads` — **Google's own reported totals** for the advertiser (from the search results page, falling back to the suggestion service), so you can tell a capped scrape from a complete one.
- `partial` — `true` on the advertiser summary when pagination ended abnormally (soft-block, stuck token, page failure) and the creative list may be incomplete; a matching `error` record explains where it was truncated.
- `format` — `image`, `html5`, or `unknown` (raw enum kept in `format_enum`).
- `image_url` — direct `simgad` image URL (static image ads).
- `preview_url` — Google `displayads-formats` preview URL (HTML5 / responsive ads).
- `first_shown` / `last_shown` — ISO-8601 UTC, converted from Google's unix timestamps.
- `regions` — populated only with `enrichCreativeDetail`: `[{geo_id, country, last_shown}]` per country.
- `variants` — populated only with `enrichCreativeDetail`: every render variant of the creative.
- `deep_link` — direct link to the creative in the Ads Transparency Center.
- `asset_kvs_key` — populated only with `downloadAssets`: the key-value-store key holding the rehosted bytes.

### Input

| Field | Type | Description |
|---|---|---|
| `advertiserName` | string | Advertiser to resolve by name (e.g. `Verizon`). Best match is scraped; other matches become `advertiser_candidate` records. Defaults to `Verizon` when no target field is provided at all. |
| `advertiserId` | string | Exact `AR…` advertiser ID. Most deterministic input; takes priority over name/domain. |
| `domain` | string | Website domain (e.g. `verizon.com`). Lowest priority. |
| `region` | string | Optional 2-letter ISO country code (US, GB, CA…; `UK`/`USA` aliased to GB/US) to filter creatives by where they ran. Empty = all regions. Any other unknown code fails fast with one error record — never an accidental all-regions scrape. |
| `maxCreativesPerAdvertiser` | integer | Cap on creatives per advertiser (`0` = no cap — full pagination). Bounds cost on very large advertisers. |
| `enrichCreativeDetail` | boolean | One extra request per creative to add per-region dates + all variants. Slower; no extra dataset items. |
| `includeAdvertiserDetail` | boolean | One request per advertiser for legal name / country / verified. Default on. |
| `downloadAssets` | boolean | Download + rehost each creative's bytes to the key-value store (URLs expire). |
| `requestDelaySeconds` | number | Politeness delay between requests (default 1s). Raise for large or enrichment-heavy crawls. |
| `proxyConfiguration` | object | Optional proxies. Works cookieless from datacenter IPs at modest volume. |

### Use cases

- **Competitive ad intelligence** — pull a competitor's entire live creative library and track how it changes week over week.
- **Creative research / swipe files** — collect every image and HTML5 ad a brand is running, by country.
- **Brand safety & compliance** — audit which ads an advertiser is currently showing in a given region.
- **Ad-tech & agency reporting** — feed an advertiser's ad inventory into dashboards with stable, rectangular fields.
- **AI agents** — a deterministic "list this advertiser's Google ads" tool with explicit input modes and clean JSON.

### FAQ

#### How do I download all of a competitor's Google ads?

Run this actor with the competitor's name, domain, or `AR…` advertiser ID. It returns every image and HTML5 creative they're running — full pagination, optional country filter, first/last-shown dates — as structured JSON you can export to CSV/Excel or feed into a dashboard. It's Google's own public transparency data, just without the one-page-at-a-time UI.

#### Is scraping the Google Ads Transparency Center legal?

The Transparency Center is Google's **own public accountability database** — the same information it shows any visitor, with no login, no gate, and no terms click-through. This actor reads that public data politely (built-in delays, no CAPTCHA/WAF evasion). See the fuller note at the bottom of this page for the nuances.

#### How does this compare to SerpApi or ad-spy tools like AdSpy and BigSpy?

SerpApi's Transparency API starts at ~$25/mo metered; ad-spy SaaS runs $9–459/mo for dashboards you can't export raw. This actor is pay-per-use (fractions of a cent per creative), returns raw structured data you own, and runs on your schedule — the right shape for automation, agencies, and AI agents rather than seat-based browsing.

#### How do I scrape Google Ads Transparency Center without an API?

Google doesn't offer a public Ads Transparency Center API. This actor talks to the same internal endpoints the website itself uses, cookieless and without a browser, and returns structured JSON — so you don't have to reverse-engineer anything.

#### How do I find an advertiser's ID?

Search by name or domain and read the `advertiser` / `advertiser_candidate` records — each carries the `AR…` ID. Or copy it from an `adstransparency.google.com/advertiser/AR…` URL. Pinning the ID is the most reliable way to run repeatedly.

#### Why do some scrapers return an empty result for a busy advertiser?

Google periodically changes the internal request payload shape. Scrapers that copied an old payload then silently receive `{}` and report zero ads. This actor mirrors the current live payload exactly and never swallows an unexpected empty response. An empty response is genuinely ambiguous, so the actor classifies it honestly: if Google's own suggestion counts say the advertiser has 0 recent and 0 lifetime ads, the zero is reported as a clean result; otherwise you get one clear `error` record stating that the advertiser may have no live ads **or** Google changed the payload / soft-blocked the request. On a region-filtered run an empty first page gets a cautious `notice` record ("zero ads in this region OR drift — cross-check without the region filter"), and an empty response mid-pagination marks the result `partial` instead of pretending it is complete.

#### Can I get all of an advertiser's ads, not just the first 100?

Yes. Leave `maxCreativesPerAdvertiser` at `0` and the actor follows Google's pagination token to the end.

#### Why are the image URLs sometimes dead when I open them later?

Google's `simgad` and preview URLs expire quickly. Enable `downloadAssets` to rehost the actual bytes into your run's key-value store, and use `asset_kvs_key`.

#### Can I filter ads by country?

Yes — set `region` to a 2-letter ISO code. The actor maps it to Google's geo ID and filters server-side.

### Pricing

This actor bills **per result** (pay-per-event, one charge per dataset item). Each creative, each advertiser summary, and each candidate record is one item. Error/notice records are capped at **one per advertiser target**, and failed requests never consume your `maxCreativesPerAdvertiser` budget — the cap counts only real, deduplicated creatives. To control cost on large advertisers, set `maxCreativesPerAdvertiser`. Enrichment flags (`enrichCreativeDetail`, `includeAdvertiserDetail`, `downloadAssets`) add requests but **not** extra dataset items.

### A note on data source and terms

This actor reads **public ad-transparency data** that Google itself publishes for accountability. The endpoints are unauthenticated and require no login or terms acceptance. Google's general Terms of Service broadly discourage automated access to its products, so treat large-scale or aggressive use as a business/legal judgment call: keep volume modest, use the built-in delays, and consider residential proxies for heavy crawls. The actor never attempts CAPTCHA or WAF evasion. You are responsible for your own use of the data.

***

**Building something on this actor?** Open an issue on the Issues tab and say what you're working on - real use cases get priority fixes and shape the roadmap.

# Actor input Schema

## `advertiserName` (type: `string`):

Advertiser to look up by name, e.g. 'Verizon' or 'Nike'. The name is resolved to an advertiser ID via the Ads Transparency Center's suggestion service; the single best match is scraped and any other matches are emitted as 'advertiser\_candidate' records so you can re-run pinned to a specific advertiserId. Ignored if advertiserId is set. Defaults to 'Verizon' (a stable, large US advertiser) when no target field is provided at all, so a run always returns data.

## `advertiserId` (type: `string`):

Exact advertiser ID starting with 'AR' (e.g. AR11385763688137883649). This is the most deterministic input — no name resolution, no ambiguity. You can find an ID from a previous run's 'advertiser' / 'advertiser\_candidate' records or from an adstransparency.google.com/advertiser/\<AR...> URL. Takes priority over advertiserName and domain.

## `domain` (type: `string`):

Look up by website domain, e.g. 'nike.com'. Resolved via the suggestion service. Note: Google's public API exposes domain match counts but NOT a domain->advertiser list, so if a domain has no directly associated advertiser candidate you will get 'domain\_candidate' records and a notice asking you to re-run with a specific advertiserId. Lowest priority (used only when advertiserId and advertiserName are both empty).

## `region` (type: `string`):

Optional ISO-3166-1 alpha-2 country code (e.g. US, GB, CA, DE) to filter creatives to those shown in that country. 'UK' and 'USA' are accepted as aliases for GB and US. Leave empty for all regions. Any other unknown code fails fast with a single error record and scrapes nothing — the run never silently falls back to an uncapped all-regions scrape.

## `maxCreativesPerAdvertiser` (type: `integer`):

Hard cap on creatives returned per advertiser (0 = no cap — follow the pagination token to the end, which is how this actor beats the incumbents' silent 100-item cap). You are billed per result, so set a cap here to bound cost for very large advertisers (some run 500+ creatives).

## `enrichCreativeDetail` (type: `boolean`):

For every creative, make one extra GetCreativeById call to add its per-region last-shown dates (regions\[]) and every render variant (variants\[]). This makes ONE additional request per creative — much slower and more likely to be rate-limited on large advertisers — but does not add extra dataset items. Leave off for fast listing.

## `includeAdvertiserDetail` (type: `boolean`):

Make one GetAdvertiserById call per advertiser to enrich the 'advertiser' summary record with legal name, country, and verified status. One cheap call per advertiser; does not add extra dataset items.

## `downloadAssets` (type: `boolean`):

Download each creative's image/preview bytes and store them in this run's key-value store, setting asset\_kvs\_key on the creative. The simgad image and preview URLs on Google's servers expire quickly, so enable this if you need the actual creative bytes to persist. Adds one download per creative.

## `requestDelaySeconds` (type: `number`):

Politeness delay between requests to adstransparency.google.com. 1 second is safe for modest volume; raise it (and add a proxy) for large or enrichment-heavy crawls to avoid soft-blocking. This actor never attempts CAPTCHA or WAF evasion.

## `proxyConfiguration` (type: `object`):

Optional proxies for the requests. When enabled, the actor rotates a 4-session proxy pool on every request — retries automatically move to a fresh IP instead of re-hitting Google from a soft-blocked one. The endpoints work cookieless from a datacenter IP at modest volume; for high-volume or enrichment-heavy runs, residential proxies plus a higher requestDelaySeconds reduce soft-blocking.

## Actor input object example

```json
{
  "advertiserName": "Verizon",
  "maxCreativesPerAdvertiser": 0,
  "enrichCreativeDetail": false,
  "includeAdvertiserDetail": true,
  "downloadAssets": false,
  "requestDelaySeconds": 1,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "advertiserName": "Verizon"
};

// Run the Actor and wait for it to finish
const run = await client.actor("silentshadow55/google-ads-transparency-actor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "advertiserName": "Verizon" }

# Run the Actor and wait for it to finish
run = client.actor("silentshadow55/google-ads-transparency-actor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "advertiserName": "Verizon"
}' |
apify call silentshadow55/google-ads-transparency-actor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=silentshadow55/google-ads-transparency-actor",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/YFIUiCtUjNA6kgvcx/builds/5VKTbDFqFMKYuhGle/openapi.json
