# Google News Scraper — Real Publisher URLs, Agent-Ready (`leorochasantos/google-news-scraper`) Actor

Fast, reliable Google News scraper: search, topic sections, and top headlines. Resolves the obfuscated Google redirect to the real publisher article URL and returns clean flat JSON built for LLM/MCP consumption.

- **URL**: https://apify.com/leorochasantos/google-news-scraper.md
- **Developed by:** [Leonardo Santos](https://apify.com/leorochasantos) (community)
- **Categories:** News, SEO tools, Business
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.30 / 1,000 article results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Google News Scraper — Real Publisher URLs, Agent-Ready

Scrape **Google News** by search query, topic section, or top headlines and get back
**clean, flat JSON** — with the one field most Google News scrapers get wrong: the
**real publisher article URL**, not Google's obfuscated redirect.

### Why this scraper

Google News RSS gives you links like
`https://news.google.com/rss/articles/CBMiigF...` — an opaque redirect, not the article.
Google changed how these are encoded in 2024, which broke a lot of scrapers (you get a
Google URL that a browser can open but your code/database/LLM can't use). **This actor
resolves each link to the actual publisher URL** (e.g. `https://www.theguardian.com/...`),
so the data is usable the moment it lands.

- ✅ **Resolved publisher URLs** — the headline feature. Falls back gracefully to the
  Google URL with `url_resolved: false` if a specific link can't be resolved, so a run
  never fails on one bad article.
- ✅ **Search, topics, and headlines** — free-text queries (with Google News operators
  like `when:1d`, `site:`, quotes), 8 topic sections, and the top-headlines feed.
- ✅ **Any language / country edition** — `hl` + `gl` (e.g. `pt-BR` / `BR`, `es-419` / `MX`).
- ✅ **Flat schema for LLMs & pipelines** — one row per article, typed fields, absolute
  URLs, explicit `url_resolved` flag. No nested junk, no formatted-string duplicates.
- ✅ **Fast & cheap** — HTTP-only, no headless browser. Runs at 512 MB.

### Input

| Field | Type | Description |
|---|---|---|
| `queries` | string\[] | Free-text Google News searches. Operators supported (`"exact phrase"`, `when:7d`, `site:bbc.com`). |
| `topics` | string\[] | Any of `WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SPORTS, SCIENCE, HEALTH`. |
| `headlines` | boolean | Include the top-headlines feed. |
| `language` | string | UI language `hl`, e.g. `en-US`, `pt-BR`. Default `en-US`. |
| `country` | string | Country edition `gl`, e.g. `US`, `BR`, `GB`. Default `US`. |
| `maxArticlesPerQuery` | integer | 1–100 per feed (RSS returns ~100 max). Default 100. |
| `resolvePublisherUrls` | boolean | Resolve real publisher URLs. Default `true`. |
| `dedupByUrl` | boolean | Drop duplicates across feeds. Default `true`. |

#### Example input

```json
{
  "queries": ["\"artificial intelligence\"", "nvidia when:1d"],
  "topics": ["TECHNOLOGY"],
  "headlines": false,
  "language": "en-US",
  "country": "US",
  "maxArticlesPerQuery": 50,
  "resolvePublisherUrls": true
}
```

### Output

One dataset item per article:

```json
{
  "query": "artificial intelligence",
  "query_type": "search",
  "language": "en-US",
  "country": "US",
  "title": "Apple sues OpenAI, alleging AI company stole trade secrets",
  "publisher": "The Guardian",
  "publisher_domain": "www.theguardian.com",
  "article_url": "https://www.theguardian.com/technology/2026/jul/10/apple-sues-openai-trade-secrets",
  "google_url": "https://news.google.com/rss/articles/CBMiigF...",
  "url_resolved": true,
  "published_at": "2026-07-10T22:33:00.000Z",
  "scraped_at": "2026-07-12T18:00:00.000Z"
}
```

### Use cases

- **News monitoring & alerts** — track a brand, ticker, or topic across publishers.
- **Media research & datasets** — build clean article-URL datasets for analysis or ML.
- **LLM / RAG pipelines** — feed resolved URLs straight into a fetch-and-summarize step.
- **Market & competitor intelligence** — topic and query feeds per country edition.

### Notes & limits

- Article **full text is not extracted** — you get the resolved URL; fetch the body
  downstream if you need it.
- Google News RSS returns up to ~100 items per feed; use multiple queries for breadth.
- Reliability is monitored continuously against a golden probe set; the target is
  **< 2% run failure** over any 30-day window.

# Actor input Schema

## `queries` (type: `array`):

Free-text Google News searches. Supports Google News operators (e.g. quotes for exact phrase, 'site:', 'when:7d'). Each query runs as its own feed.

## `topics` (type: `array`):

Google News topic sections to pull top articles from. Leave empty if you only want search or headlines.

## `headlines` (type: `boolean`):

Also include the main top-headlines feed for the chosen language/country.

## `language` (type: `string`):

UI language code, e.g. 'en-US', 'pt-BR', 'es-ES'. Controls the language edition of the feed.

## `country` (type: `string`):

Two-letter country edition, e.g. 'US', 'BR', 'GB'. Controls which country's Google News edition is queried.

## `maxArticlesPerQuery` (type: `integer`):

Cap on articles taken from each feed (1–100). Google News RSS returns up to ~100 items per feed.

## `resolvePublisherUrls` (type: `boolean`):

Resolve each obfuscated Google redirect to the actual publisher article URL. This is the actor's headline feature. Disable for a faster run if the Google redirect URL is enough for you.

## `dedupByUrl` (type: `boolean`):

Drop duplicate articles that appear in more than one feed (matched by resolved URL when available).

## `proxyConfiguration` (type: `object`):

Apify Proxy recommended for reliable throughput. Defaults to the automatic Apify Proxy group.

## Actor input object example

```json
{
  "queries": [
    "\"climate change\"",
    "bitcoin when:1d"
  ],
  "topics": [],
  "headlines": false,
  "language": "en-US",
  "country": "US",
  "maxArticlesPerQuery": 100,
  "resolvePublisherUrls": true,
  "dedupByUrl": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "artificial intelligence"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("leorochasantos/google-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": ["artificial intelligence"] }

# Run the Actor and wait for it to finish
run = client.actor("leorochasantos/google-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "artificial intelligence"
  ]
}' |
apify call leorochasantos/google-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=leorochasantos/google-news-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/Aj1LXtObWGsTccLfs/builds/TOo78QxxVcNbZhPRa/openapi.json
