# GDELT News Scraper - Worldwide Coverage, No API Key (`dami_studio/gdelt-news-scraper`) Actor

Monitor what the whole world is publishing about your topic. Searches the public GDELT 2.0 global news index and returns clean rows: headline, URL, outlet domain, source country, language, publish time and social image. Filter by keyword, timespan, country and language. No API key, no quota.

- **URL**: https://apify.com/dami\_studio/gdelt-news-scraper.md
- **Developed by:** [Dami's Studio](https://apify.com/dami_studio) (community)
- **Categories:** News, AI, Integrations
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

$2.00 / 1,000 article returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## GDELT Worldwide News Scraper

Search worldwide news through the [GDELT 2.0 DOC API](https://blog.gdeltproject.org/gdelt-doc-2-0-api-debuts/) — a public, no-key, no-login JSON endpoint that indexes online news from across the planet in 65+ languages. No proxy or anti-bot handling needed.

### What it does

Given a search query, the actor calls the GDELT DOC API in `ArtList` mode and returns clean, structured articles. You can filter by recency, source country, and source language, and sort by newest, oldest, or relevance.

### Input

| Field | Type | Default | Notes |
|---|---|---|---|
| `query` | string | — (required) | Keywords. Quote phrases: `"climate change"`. Very short/common single words may be rejected by GDELT. |
| `maxItems` | integer | `100` | Capped at **250** — GDELT returns at most 250 articles per request. |
| `sort` | string | `DateDesc` | `DateDesc` (newest), `DateAsc` (oldest), `HybridRel` (relevance). |
| `timespan` | string | — | Recency window, e.g. `1d`, `3d`, `1w`, `1m`, `3m`. GDELT covers ~the last 3 months. |
| `sourceCountry` | string | — | Appended as `sourcecountry:{code}` (e.g. `US`, `UK`, `FR`). |
| `sourceLang` | string | — | Appended as `sourcelang:{code}` (e.g. `english`, `french`). |
| `proxyConfiguration` | object | — | Optional; not needed (public API). |

### Output

Each successful row:

```json
{
  "ok": true,
  "title": "…",
  "url": "https://…",
  "domain": "example.com",
  "sourceCountry": "United States",
  "language": "English",
  "publishedAt": "2026-06-11T12:00:00.000Z",
  "socialImage": "https://…"
}
```

Results are de-duplicated by URL. Each `ok:true` article is billed one `article` charge unit. Diagnostic rows (`ok:false`) and empty/blocked runs are never charged.

**Nullable fields:** GDELT does not always populate every field. Any of `title`, `url`, `domain`, `sourceCountry`, `language`, `publishedAt`, and `socialImage` can be `null` for a given article (e.g. `socialImage` is often missing, and `publishedAt` is `null` when GDELT's `seendate` is unparseable). Rows with neither a `url` nor a `title` are dropped before charging.

### Diagnostics

The actor never fails silently. Instead it writes a single diagnostic row (`ok:false`) with an `errorCode` and never charges for it:

- `BAD_INPUT` — GDELT rejected the query (e.g. "query too short"). Quote phrases and avoid overly short/common terms.
- `NO_RESULTS` — the query was valid but matched nothing. Broaden it or widen the timespan.
- `RATE_LIMITED` / `SERVER_ERROR` / `NETWORK` — transient issues; the actor retried with backoff first.

### Notes / quirks

- GDELT requires the query to be URL-encoded and phrases to be quoted — the actor handles both.
- On a malformed query GDELT may return a `text/plain` error string (sometimes with HTTP 200) or an empty body instead of JSON. The actor guards `JSON.parse` and surfaces a clear `BAD_INPUT` diagnostic.
- GDELT's index covers roughly the last 3 months of news.
- The actor rotates a real browser **User-Agent** per request attempt for retry resilience, and supports an **optional proxy** (`proxyConfiguration`). Neither is required — GDELT is a public no-key API with no anti-bot — so leave the proxy unset unless you hit IP-level rate limits.

# Actor input Schema

## `query` (type: `string`):

Keywords to search worldwide news for. Wrap phrases in quotes (e.g. "climate change"). Combine with OR, or operators like domain:reuters.com. Very short or very common single words may be rejected by GDELT — add a second word or quote a phrase.

## `maxItems` (type: `integer`):

Maximum number of articles to return. GDELT hard-caps a single request at 250 — higher values are automatically capped at 250.

## `sort` (type: `string`):

How to order results: DateDesc (newest first), DateAsc (oldest first), or HybridRel (by relevance to the query).

## `timespan` (type: `string`):

Only return articles from the last N units of time, e.g. 1d, 3d, 1w, 1m, 3m. Leave empty for GDELT's default window. GDELT only covers roughly the last 3 months.

## `sourceCountry` (type: `string`):

Limit to articles from a given source country, appended to the query as sourcecountry:{code}. Use GDELT country codes (e.g. US, UK, FR, DE, IN). Leave empty for all countries.

## `sourceLang` (type: `string`):

Limit to articles in a given language, appended to the query as sourcelang:{code} (e.g. english, french, spanish, german). Leave empty for all languages.

## `notionConnector` (type: `string`):

Optional. Write each article as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default) — results are always saved to the dataset regardless.

## `notionParentId` (type: `string`):

Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead.

## `proxyConfiguration` (type: `object`):

Optional proxy for the GDELT requests. GDELT is a public no-key API with no anti-bot, so a proxy is not needed in most cases. Leave unset to connect directly.

## Actor input object example

```json
{
  "query": "artificial intelligence",
  "maxItems": 100,
  "sort": "DateDesc",
  "timespan": "",
  "sourceCountry": "",
  "sourceLang": ""
}
```

# Actor output Schema

## `results` (type: `string`):

Scraped rows are stored in the default dataset (one row per result). Blocked/empty/error runs return a single uncharged diagnostic row instead.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "artificial intelligence"
};

// Run the Actor and wait for it to finish
const run = await client.actor("dami_studio/gdelt-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "artificial intelligence" }

# Run the Actor and wait for it to finish
run = client.actor("dami_studio/gdelt-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "artificial intelligence"
}' |
apify call dami_studio/gdelt-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=dami_studio/gdelt-news-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/5LUPX8Wm6SBLk7wZe/builds/DNaSPv7t92bHRsF30/openapi.json
