# arXiv Research Paper Scraper (`seeb/arxiv-research-paper-scraper`) Actor

Scrape arXiv papers by keyword or category and return research titles, abstracts, authors, dates, links, and topic signals.

- **URL**: https://apify.com/seeb/arxiv-research-paper-scraper.md
- **Developed by:** [Techionik](https://apify.com/seeb) (community)
- **Categories:** News, Developer tools, Automation
- **Stats:** 1 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.20 / 1,000 arxiv paper rows

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## arXiv Research Paper Scraper

Find arXiv papers for AI research change detection, academic trend tracking, technical due diligence, startup research, and market intelligence.

This Actor uses the public arXiv API. It is fast, low-cost, and does not need a proxy by default.

### What It Returns

- Paper title, arXiv URL, and PDF URL
- Authors and categories
- Published and updated timestamps
- Source query or source category
- Relevance score and matched source terms
- Optional abstract text
- Change Detection mode for newly discovered papers

### Common Uses

- Track new AI/ML papers by topic
- Scraper research categories such as `cs.AI`, `cs.LG`, `cs.CL`, and `cs.CV`
- Build technical due diligence datasets
- Watch emerging academic trends
- Research startup/market themes
- Create alerts for new papers around a keyword

### Input Tips

Use focused research phrases:

- `large language model agents`
- `retrieval augmented generation`
- `computer vision transformer`
- `diffusion model`
- `robot learning`
- `graph neural network`

Use `maxResultsPerSource` when you provide several queries/categories and want balanced output.

### Result Limits

`maxResults` is a maximum cap, not a guarantee. The Actor only writes papers that pass your relevance filters.

The minimum `maxResults` is 50 so runs are commercially practical and users are charged for meaningful output.

### Change Detection

Enable `monitorChanges` to compare the current run against the previous snapshot in the Actor key-value store.

When `onlyChanges` is enabled, the dataset contains only newly detected papers. When it is disabled, the dataset contains the current full result set while the change snapshot is still stored for automation.

### Proxy Use

No proxy is needed by default. This Actor uses the public arXiv API and does not run a browser.

### Output

Each dataset row represents one arXiv paper. The primary output event is one dataset row.

### Notes

This Actor is not affiliated with arXiv or Cornell University. It reads public arXiv API results and stores structured rows in your Apify dataset.

# Actor input Schema

## `searchQueries` (type: `array`):

Research search terms. Use AI, ML, robotics, biotech, finance, security, or technical market phrases.

## `categories` (type: `array`):

Optional arXiv categories such as cs.AI, cs.LG, cs.CL, cs.CV, stat.ML, q-fin, or eess.

## `sortBy` (type: `string`):

Submitted date is best for change detection fresh papers.

## `sortOrder` (type: `string`):

Descending returns newest or most relevant results first.

## `maxResults` (type: `integer`):

Maximum paper rows to write for the whole run.

## `maxResultsPerSource` (type: `integer`):

Optional cap per query/category so one broad source does not fill the entire run. Use 0 for no source cap.

## `pagesPerSource` (type: `integer`):

How many arXiv API pages to scan per source. Each page requests up to 100 papers.

## `includeAbstract` (type: `boolean`):

Include paper abstract/summary in each dataset row.

## `minimumRelevanceScore` (type: `integer`):

Minimum number of meaningful source words that must appear in the title, abstract, category, or author list.

## `monitorChanges` (type: `boolean`):

Compare against the previous saved snapshot and detect newly found papers.

## `onlyChanges` (type: `boolean`):

When change detection is enabled, write only newly detected papers to the dataset.

## Actor input object example

```json
{
  "searchQueries": [
    "diffusion model",
    "robot learning",
    "graph neural network"
  ],
  "categories": [
    "cs.AI",
    "cs.LG",
    "cs.CL"
  ],
  "sortBy": "submittedDate",
  "sortOrder": "descending",
  "maxResults": 50,
  "maxResultsPerSource": 0,
  "pagesPerSource": 1,
  "includeAbstract": true,
  "minimumRelevanceScore": 0,
  "monitorChanges": false,
  "onlyChanges": false
}
```

# Actor output Schema

## `results` (type: `string`):

Normal runs contain arXiv paper rows. When output-only-changes is enabled, rows contain newly detected papers.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "machine learning"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("seeb/arxiv-research-paper-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchQueries": ["machine learning"] }

# Run the Actor and wait for it to finish
run = client.actor("seeb/arxiv-research-paper-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "machine learning"
  ]
}' |
apify call seeb/arxiv-research-paper-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=seeb/arxiv-research-paper-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/0lMe22AYDejQdzj8f/builds/5iQ0zJvuQ2d35LpVP/openapi.json
