# Academic Papers Scraper — OpenAlex, 250M+ Works, No Login (`charliemorrisondev/openalex-papers-scraper`) Actor

Search 250M+ scholarly works from OpenAlex — by keyword, author, institution or journal; filter by year, citations, open access and work type. No login, no API key, no captcha. Normalized JSON: DOI, title, authors, venue, year, citations, OA PDF link, abstract.

- **URL**: https://apify.com/charliemorrisondev/openalex-papers-scraper.md
- **Developed by:** [Petro Pankov](https://apify.com/charliemorrisondev) (community)
- **Categories:** Education, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 papers

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Academic & Research Papers Scraper — OpenAlex (250M+ works, no login)

Search the world's scholarly literature — research papers, journal articles, preprints,
book chapters, datasets and reviews — and get clean, structured rows back. Built directly on
[OpenAlex](https://openalex.org), an open catalogue of over 250 million works.

**No login. No API key. No captcha. No proxies needed.**

### Why this instead of a Google Scholar scraper

Google Scholar has no public API. Every scraper built on it is reverse-engineering an
HTML surface that actively fights automation — which is why those Actors break, rate-limit,
and need proxy budgets. OpenAlex publishes the same literature through a documented, open,
versioned API with a stable schema. Nothing to fight, nothing to break.

That means: predictable runs, no proxy costs, no captcha failures, and results you can
actually schedule against.

### What you can do

- **Keyword search** across title, abstract and fulltext — `"large language models"`
- **By author** — `"Yoshua Bengio"` (name match, no author IDs needed)
- **By institution** — `"Massachusetts Institute of Technology"`, `"Max Planck"`
- **By journal / conference** — `"Nature"`, `"NeurIPS"`
- **Look up by DOI & pull citation counts** — every result carries a bare `doi` and a
  `citedByCount`, so you can resolve DOIs to metadata or rank papers by citations
- **Filter** by year range, minimum citations, open-access-only, and work type
- **Sort** by relevance, citation count, or newest first

### Output

Each row is normalized:

| field | description |
|---|---|
| `id`, `openAlexUrl` | OpenAlex work ID and URL |
| `doi` | DOI (bare, no `https://doi.org/` prefix) |
| `title` | Work title |
| `authors`, `firstAuthor`, `authorCount` | Full author list and convenience fields |
| `venue`, `publisher`, `issn` | Journal / conference and publisher |
| `year`, `publishedAt` | Publication year and ISO date |
| `type`, `language` | Work type (article, preprint, …) and language |
| `citedByCount`, `referencedWorksCount`, `fwci` | Citation metrics |
| `isOpenAccess`, `oaStatus`, `pdfUrl` | Open-access status and direct free-PDF link |
| `landingPageUrl` | Publisher landing page |
| `concepts`, `institutions`, `countries` | Topics and affiliations |
| `abstract` | Full abstract text (when **Include abstract** is on) |

### Example

```json
{
  "search": "retrieval augmented generation",
  "fromYear": 2023,
  "minCitations": 10,
  "openAccessOnly": true,
  "sortBy": "citations",
  "maxResults": 200,
  "includeAbstract": true
}
```

### Typical uses

- Literature reviews and systematic reviews
- Citation analysis and citation-count tracking across a field or author
- DOI enrichment — resolve a list of DOIs into full titles, authors, venues and metrics
- Tracking a lab's, author's or institution's output
- Building datasets of open-access PDFs for downstream NLP
- Competitive/landscape research in a scientific field

### FAQ

**Is this a Google Scholar alternative?** Yes — it covers the same research-paper literature
(and more) through OpenAlex's open API, without Google Scholar's captchas, rate limits or proxy
costs. See the comparison above.

**Can I look up papers by DOI or get citation counts?** Every row carries the `doi`, `citedByCount`
and `referencedWorksCount` fields, so you can pull citations, references and DOIs for any matched work.

**Does it export research papers as a dataset?** Yes — results come back as a flat dataset (JSON,
CSV, Excel) ready for literature reviews, citation analysis, or building NLP/RAG corpora.

**Can I use it for bibliometrics?** Yes — each row carries citation counts, referenced-works counts
and OpenAlex's field-weighted citation impact (`fwci`), so you can run bibliometric and
research-impact analysis without a separate metrics source.

**Does it cover biomedical / PubMed papers?** OpenAlex indexes every discipline, including
biomedical and PubMed-indexed literature, so medical and life-science works show up in the same search.

**Can I get only open access papers?** Yes — set `openAccessOnly` and every row comes back with
`isOpenAccess`, the `oaStatus` (gold, green, hybrid, bronze) and a direct `pdfUrl` where a free
full text exists, so you can build a clean list of open-access papers with downloadable PDFs.

**Can I scrape preprints?** Yes — preprints are a work `type` in OpenAlex, so you can pull them
with the rest of the literature or restrict a run to preprints via the `types` filter (e.g. arXiv,
bioRxiv and other preprint servers indexed by OpenAlex).

**Can I use it for a literature review?** Yes — it searches the world's scientific literature by
keyword, author, journal and year, so you can assemble the full set of relevant papers for a
literature review in one run and export them with abstracts, DOIs and citation counts attached.

**Which databases does it replace?** It taps the same corpus that powers tools built on Crossref,
Microsoft Academic Graph, PubMed and Semantic Scholar, unified into one open source.

### Notes

- Results are capped at today's date unless you set **To year** — OpenAlex contains a small
  number of records with future publication dates, and this keeps "newest first" meaningful.
- Requests go through OpenAlex's polite pool (a contact address is sent with each call), which
  is the usage pattern OpenAlex asks for.
- Abstracts are stored by OpenAlex as an inverted index for licensing reasons; this Actor
  reconstructs them into plain text for you.

# Actor input Schema

## `search` (type: `string`):

Full-text search across title, abstract and fulltext (e.g. "large language models", "CRISPR off-target"). Leave empty if you are filtering by author, institution, venue or year instead.

## `authors` (type: `array`):

Match works by author name, e.g. "Yoshua Bengio". Multiple names are OR-ed.

## `institutions` (type: `array`):

Match works by author affiliation, e.g. "MIT", "Max Planck". Multiple values are OR-ed.

## `venues` (type: `array`):

Match works by journal or conference name, e.g. "Nature", "NeurIPS". Multiple values are OR-ed.

## `fromYear` (type: `integer`):

Only works published in this year or later.

## `toYear` (type: `integer`):

Only works published in this year or earlier. If omitted, results are capped at today — OpenAlex contains some records with bogus future publication dates.

## `minCitations` (type: `integer`):

Only works cited at least this many times. Useful for cutting long-tail noise.

## `openAccessOnly` (type: `boolean`):

Only works with a free full text available.

## `types` (type: `array`):

Restrict to work types, e.g. article, preprint, book-chapter, dataset, review.

## `sortBy` (type: `string`):

relevance (default, requires a search query), citations (most cited first), or date (newest first).

## `maxResults` (type: `integer`):

Maximum number of works to return. Results are paged 200 at a time.

## `includeAbstract` (type: `boolean`):

Reconstruct and include the full abstract text for each work. Slightly larger dataset rows.

## Actor input object example

```json
{
  "search": "large language models",
  "authors": [
    "Yoshua Bengio"
  ],
  "institutions": [
    "Massachusetts Institute of Technology"
  ],
  "venues": [
    "Nature"
  ],
  "fromYear": 2023,
  "minCitations": 10,
  "openAccessOnly": false,
  "types": [
    "article"
  ],
  "sortBy": "relevance",
  "maxResults": 100,
  "includeAbstract": false
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "search": "large language models"
};

// Run the Actor and wait for it to finish
const run = await client.actor("charliemorrisondev/openalex-papers-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "search": "large language models" }

# Run the Actor and wait for it to finish
run = client.actor("charliemorrisondev/openalex-papers-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "search": "large language models"
}' |
apify call charliemorrisondev/openalex-papers-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=charliemorrisondev/openalex-papers-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/q8nVTyAVkFJcOjTvI/builds/FRxHZ6WwzFSfmHP09/openapi.json
