# News Article & RSS Crawler — Clean Text for RAG (`ahampton83/news-article-crawler`) Actor

Fetch news from RSS feeds and Google News search, then extract clean article text and metadata. Perfect for RAG pipelines, newsletters, trend monitoring, and AI agents. Use via Apify Console/API or connect as an MCP server for Claude, Cursor, and other AI agents.

- **URL**: https://apify.com/ahampton83/news-article-crawler.md
- **Developed by:** [Aaron Hampton](https://apify.com/ahampton83) (community)
- **Categories:** News, AI, MCP servers
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 article fetcheds

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## News Article & RSS Crawler

Fetch news from **RSS/Atom feeds** and **Google News search**, then extract clean article text with metadata. Works as a normal Apify Actor or as an **MCP server** for Claude, Cursor, and other AI agents.

### Features

- **Google News search** via Google's public RSS endpoint
- **Direct RSS/Atom feeds** — any standard feed
- **Clean article extraction** — strips nav, ads, footers, and boilerplate, returning readable body text
- **MCP Server** — `search_news`, `fetch_feed`, `get_article_text` tools
- **Pay-per-event pricing** — first 20 articles free per run/tool call

### MCP Tools

| Tool | Description |
|------|-------------|
| `search_news(query, maxResults?, includeFullText?, geo?)` | Search Google News and extract articles |
| `fetch_feed(feedUrls, maxItems?, includeFullText?)` | Fetch articles from RSS/Atom feeds |
| `get_article_text(url)` | Extract clean text from a single article URL |

### Normal Mode Input

```json
{
    "query": "artificial intelligence",
    "maxArticles": 20,
    "includeFullText": true,
    "geo": "US"
}
```

Or pass direct feeds instead of a query:

```json
{
    "feedUrls": ["https://feeds.bbci.co.uk/news/rss.xml"],
    "maxArticles": 10,
    "includeFullText": true
}
```

### Output

```json
{
    "query": "artificial intelligence",
    "feedUrls": [],
    "totalResults": 20,
    "errors": [],
    "articles": [
        {
            "title": "AI Agents Reach New Milestone",
            "url": "https://example.com/ai-agents",
            "resolvedUrl": "https://example.com/ai-agents",
            "source": "Example Source",
            "publishedAt": "2026-07-04T12:00:00.000Z",
            "summary": "...",
            "text": "Clean, extracted article body text...",
            "wordCount": 842,
            "fetchedAt": "2026-07-04T19:00:00.000Z"
        }
    ],
    "fetchedAt": "2026-07-04T19:00:00.000Z"
}
```

### Pricing (Pay-Per-Event)

| Event | Price | Description |
|-------|-------|-------------|
| Actor start | $0.00005 | Billed once per run |
| Article fetched | $0.005 | Per article successfully extracted (first 20 free per run/tool call) |
| MCP tool call | $0.005 | Per MCP tool invocation |

### Technical Approach

The Actor uses `got-scraping` with optional Apify proxy rotation for HTTP requests. RSS/Atom XML is parsed with `fast-xml-parser`. Article text is extracted with a Cheerio heuristic that removes scripts, navigation, sidebars, ads, and footers, then picks the richest `<article>`, `<main>`, or content-class container.

### Development

```bash
npm install        # Install dependencies
npm run build      # Compile TypeScript
npm test           # Run tests
npm run start:dev  # Run locally (development mode)
```

### Categories

`NEWS`, `AI`, `MCP_SERVERS`

### License

ISC

# Actor input Schema

## `query` (type: `string`):

Search term for Google News RSS (e.g. 'artificial intelligence'). Either query or feedUrls must be provided.

## `feedUrls` (type: `array`):

One or more direct RSS or Atom feed URLs to fetch.

## `maxArticles` (type: `integer`):

Maximum number of articles to return.

## `includeFullText` (type: `boolean`):

Fetch and extract the full article body for each item. Disable to return only headlines and metadata.

## `geo` (type: `string`):

ISO country code for Google News region (e.g. US, GB, DE). Defaults to US.

## Actor input object example

```json
{
  "query": "",
  "feedUrls": [],
  "maxArticles": 20,
  "includeFullText": true,
  "geo": "US"
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("ahampton83/news-article-crawler").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("ahampton83/news-article-crawler").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call ahampton83/news-article-crawler --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=ahampton83/news-article-crawler",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/HM986YN1bM8uG7se5/builds/40kbqxqAz5RVxeOgr/openapi.json
