# arXiv Research Papers & Abstracts Scraper (`scrapers_lat/arxiv-papers-scraper`) Actor

Scrape arXiv preprints by keyword, author or subject with arXiv ID, title, authors, abstract, subject categories, DOI, publication and update dates and PDF links. Export to JSON, CSV or Excel.

- **URL**: https://apify.com/scrapers\_lat/arxiv-papers-scraper.md
- **Developed by:** [Scrapers Lat](https://apify.com/scrapers_lat) (community)
- **Categories:** Developer tools, News, AI
- **Stats:** 2 total users, 1 monthly users, 90.9% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $6.80 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## arXiv Research Papers & Abstracts Scraper

> Search arXiv and export clean, structured paper metadata: arXiv ID, title, authors, full abstract, subject categories, DOI, publication and update dates and PDF links. Perfect for literature reviews, research dashboards and citation tracking.

**📥 [Input](https://apify.com/scrapers_lat/arxiv-papers-scraper/input-schema) · 📤 [Output](https://apify.com/scrapers_lat/arxiv-papers-scraper/output-schema) · 💰 [Pricing](https://apify.com/scrapers_lat/arxiv-papers-scraper/pricing) · ▶️ [Examples](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples)**

![Apify](https://img.shields.io/badge/Platform-Apify-1CE1CE?logo=apify\&logoColor=white)
![Research papers](https://img.shields.io/badge/Data-Research%20papers-blue)
![Output](https://img.shields.io/badge/Output-JSON%20%7C%20CSV%20%7C%20Excel-orange)

<table><tr>
<td align="center"><strong>Full abstracts</strong><br>plus authors</td>
<td align="center"><strong>Categories & DOI</strong><br>with PDF links</td>
<td align="center"><strong>JSON / CSV / Excel</strong><br>output formats</td>
</tr></table>

<br>

### Who is it for

| Use case | Who benefits |
|---|---|
| Literature reviews | Researchers building a corpus on a topic fast |
| Research dashboards | Teams tracking new work in a field |
| Citation and trend analysis | Analysts studying authors, categories and volume over time |
| ML datasets | Builders assembling training or evaluation sets of abstracts |

### How to use it

1. Enter a **search query** (for example `large language models`, `quantum computing` or `protein folding`).
2. Optionally choose which **field** to match (title, abstract, author or category) and how to **sort** (relevance, newest, recently updated).
3. Set **Max Items** and run. Export as JSON, CSV or Excel, or pull it through the Apify API.

### Frequently Asked Questions

**Can I search by author or subject category?**
Yes. Set the search field to Author or Category code, or pass a native query such as `au:hinton` or `cat:cs.CL`. You can also combine terms with AND / OR.

**Do I get the full abstract?**
Yes. Each record includes the complete abstract text, not just a snippet.

**Does every paper have a DOI?**
No. A DOI is included whenever the paper has one registered; otherwise the field is null. Every record still has a stable arXiv ID and links.

**How many papers can I collect?**
Set Max Items to whatever you need. Results are gathered page by page until that limit or the end of the matches is reached.

### Example use cases

Ready-to-run example tasks, each preconfigured for a common scenario. Open one and press run, or use it as a template:

- [Large Language Models research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-llm-papers-relevance): Search arXiv for large language models research papers with titles, authors, abstracts, categories, dates, and PDF links.
- [Diffusion Models research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-diffusion-models-newest): Search arXiv for diffusion models research papers with titles, authors, abstracts, categories, dates, and PDF download links.
- [Reinforcement Learning research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-reinforcement-learning-updated): Search arXiv for reinforcement learning research papers with titles, authors, abstracts, categories, dates, and PDF links.
- [Graph Neural Networks research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-graph-neural-networks-newest): Search arXiv for graph neural networks research papers with titles, authors, abstracts, categories, dates, and PDF links.
- [Quantum Computing research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-quantum-computing-relevance): Search arXiv for quantum computing research papers with titles, authors, abstracts, categories, publication dates, and PDF links.
- [Protein Folding research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-protein-folding-newest): Search arXiv for protein folding research papers with titles, authors, abstracts, categories, publication dates, and PDF links.
- [Computer Vision Transformers research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-vision-transformers-relevance): Search arXiv for computer vision transformer papers with titles, authors, abstracts, categories, dates, and PDF download links.
- [Federated Learning research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-federated-learning-updated): Search arXiv for federated learning research papers with titles, authors, abstracts, categories, dates, and PDF download links.
- [Computation and Language papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-cs-cl-category-newest): Search arXiv computation and language papers in cs.CL with titles, authors, abstracts, categories, dates, and PDF links.
- [Research papers by Yoshua Bengio from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-author-yoshua-bengio-newest): Search arXiv for papers by Yoshua Bengio with titles, co-authors, abstracts, categories, publication dates, and PDF links.

### Export, API and AI agents (x402 + MCP)

Export the scraped data to **JSON, CSV or Excel**, pull it as a **dataset** through the Apify **API**, or wire it into your app with **no code**. This web scraper and data extractor also works for bulk data extraction and scheduled runs.

For AI agents: this Actor is available on **x402**, Apify's agentic payment standard built with Coinbase. An AI agent can discover, pay for and run it on its own with a funded wallet and a single HTTP request: no account, no subscription, no API key and no human in the loop. It also runs as an **MCP** tool inside Claude, Cursor and other AI clients out of the box. Learn more about [x402 agentic payments on Apify](https://docs.apify.com/platform/integrations/x402).

### Related scrapers

- [Crossref Works Scraper](https://apify.com/scrapers_lat/crossref-scraper)
- [Google News Scraper](https://apify.com/scrapers_lat/google-news-scraper)

### More scrapers at scrapers.lat

This actor is built and maintained by [scrapers.lat](https://scrapers.lat), where we publish scrapers for public platforms: finance, news, real estate, jobs, e-commerce and government data. Browse the full catalog or ask us for a custom scraper at [scrapers.lat](https://scrapers.lat).

***

> This actor is an independent tool and has no affiliation with arXiv or Cornell University. It only accesses publicly available paper metadata. Use the results in accordance with the source's terms.

# Actor input Schema

## `maxPapers` (type: `integer`):

Maximum number of papers to collect. Optional.

## `searchQuery` (type: `string`):

Words or phrase to search across arXiv papers (for example 'large language models', 'quantum computing', 'protein folding'). Advanced users can pass a native arXiv query with field prefixes such as 'ti:transformer AND cat:cs.CL'.

## `field` (type: `string`):

Which field to match the search query against when a plain phrase is entered.

## `sortBy` (type: `string`):

Order the results by relevance, newest submitted, or most recently updated.

## Actor input object example

```json
{
  "maxPapers": 50,
  "searchQuery": "large language models",
  "field": "all",
  "sortBy": "relevance"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxPapers": 50,
    "searchQuery": "large language models"
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapers_lat/arxiv-papers-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxPapers": 50,
    "searchQuery": "large language models",
}

# Run the Actor and wait for it to finish
run = client.actor("scrapers_lat/arxiv-papers-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxPapers": 50,
  "searchQuery": "large language models"
}' |
apify call scrapers_lat/arxiv-papers-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scrapers_lat/arxiv-papers-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/pVgQ8Av6QjxF99wRz/builds/PcHH9joMFVnangM67/openapi.json
