# arXiv Papers Scraper: AI & Science Research Tracker (`scrapemint/arxiv-papers-scraper`) Actor

Track new research papers on arXiv by keyword, category, or author. One clean JSON row per paper: title, abstract, authors, categories, dates, PDF link, and DOI. Official open API, no key, no browser. Pay per paper.

- **URL**: https://apify.com/scrapemint/arxiv-papers-scraper.md
- **Developed by:** [Ken M](https://apify.com/scrapemint) (community)
- **Categories:** AI, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$3.00 / 1,000 arxiv paper rows

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## arXiv Papers Scraper: AI & Science Research Tracker

Track new research the moment it posts. Give this Actor keywords, arXiv categories (cs.AI, cs.LG, cs.CL, stat.ML, or any other), or author names and it returns one clean JSON row per paper: title, full abstract, authors, categories, submitted and updated dates, PDF link, DOI, and journal reference. It reads arXiv's official open API, so there is no key, no browser, no proxy, and no rate-limit games, ever.

Built for AI teams tracking specific topics, newsletter and content operators, analysts watching research trends, and RAG builders assembling paper corpora. Turn on cross-run dedupe, put it on a daily schedule, and each run returns only the papers you have not seen yet.

### What you get

One row per paper, with:

- `title`, `abstract` (full text), `authors`
- `primaryCategory`, `categories` (cross-lists included)
- `published`, `updated`, `doi`, `journalRef`, `comment` (venue/pages notes)
- `arxivId`, `url` (abstract page), `pdfUrl`

### Input

- `searchQueries` (keywords matched on title + abstract, OR-combined)
- `categories` (arXiv category codes, OR-combined; AND-ed with keywords)
- `authors` (optional author filter)
- `dateFrom` (only papers on/after this date)
- `sortBy` (newest submitted, recently updated, or relevance)
- `maxPapers` (default 25, up to 1000)
- `dedupe` (skip papers returned by previous runs; built for scheduled monitoring)

### Example input

```json
{
  "searchQueries": ["multi-agent", "tool use"],
  "categories": ["cs.AI", "cs.CL"],
  "dateFrom": "2026-07-01",
  "maxPapers": 50,
  "dedupe": true
}
```

### Example output

```json
{
  "arxivId": "2507.01234v1",
  "url": "https://arxiv.org/abs/2507.01234v1",
  "pdfUrl": "https://arxiv.org/pdf/2507.01234v1",
  "title": "Distributed Attacks in Persistent-State AI Control",
  "abstract": "We study control protocols for AI agents that maintain persistent state...",
  "authors": ["Jane Doe", "John Smith"],
  "primaryCategory": "cs.AI",
  "categories": ["cs.AI", "cs.CR"],
  "published": "2026-07-02T17:58:01Z",
  "updated": "2026-07-02T17:58:01Z",
  "doi": null,
  "journalRef": null
}
```

### Uses

- Daily digest of new papers in your niche, deduped, straight into Slack or a newsletter draft
- Track what a specific lab or author publishes
- Build topic-scoped paper corpora (abstracts included) for RAG and embeddings
- Watch research momentum on a technology before the market does
- Feed a literature review with structured rows instead of browser tabs

### Pricing

Pay per paper. Queries that match nothing cost nothing. The first 2 rows of every run are free so you can validate output before you scale up.

### Notes

- Uses arXiv's official public API, which is intended for programmatic access. The Actor honors the API's politeness guidance (about one request every 3 seconds when paginating).
- arXiv covers physics, math, CS, biology, finance, statistics, EE, and economics; category codes are on arxiv.org/category\_taxonomy.

# Actor input Schema

## `searchQueries` (type: `array`):

Keywords matched against title and abstract (OR-combined). Example: multi-agent, retrieval augmented generation. Leave empty to browse categories only.

## `categories` (type: `array`):

Category codes (OR-combined), e.g. cs.AI, cs.LG, cs.CL, stat.ML. Combined with keywords using AND.

## `authors` (type: `array`):

Author names to filter by (OR-combined), e.g. Yoshua Bengio.

## `dateFrom` (type: `string`):

ISO date, e.g. 2026-07-01. Filters on the submitted date (or updated date when sorting by last update).

## `sortBy` (type: `string`):

Newest submissions first is the monitoring default.

## `maxPapers` (type: `integer`):

Cap on papers returned. Controls total cost.

## `dedupe` (type: `boolean`):

Remember returned arXiv IDs across runs and skip them. Turn on for scheduled monitoring so each run returns only new papers.

## Actor input object example

```json
{
  "searchQueries": [
    "multi-agent"
  ],
  "categories": [
    "cs.AI"
  ],
  "authors": [],
  "dateFrom": "",
  "sortBy": "submittedDate",
  "maxPapers": 25,
  "dedupe": false
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "multi-agent"
    ],
    "categories": [
        "cs.AI"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapemint/arxiv-papers-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": ["multi-agent"],
    "categories": ["cs.AI"],
}

# Run the Actor and wait for it to finish
run = client.actor("scrapemint/arxiv-papers-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "multi-agent"
  ],
  "categories": [
    "cs.AI"
  ]
}' |
apify call scrapemint/arxiv-papers-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scrapemint/arxiv-papers-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/eZgfLEeQG5AJRWHuK/builds/1h0C9GyKxEqq1cipk/openapi.json
