# Crossref Scraper — Scholarly Works, DOIs & Citations (`hipersoft/crossref-scraper`) Actor

Bulk-scrape scholarly works from Crossref's 150M+ record index: DOI, title, authors (with ORCID), journal, publisher, publication date, citation count, references, ISSN, subjects and abstract. Full-text search, date and type filters, thousands of results per run. No login, no API key.

- **URL**: https://apify.com/hipersoft/crossref-scraper.md
- **Developed by:** [hiper soft](https://apify.com/hipersoft) (community)
- **Categories:** Other
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.0004 / work scraped

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Crossref Scraper — Bulk Scholarly Works, DOIs, Citations & Metadata

Bulk-scrape scholarly metadata from **Crossref**, the DOI registry indexing **150M+ works** across virtually every academic publisher. Get DOIs, titles, authors (with ORCID), journals, publishers, publication dates, **citation counts**, references, ISSNs, subjects and abstracts. Full-text search plus date and type filters. Thousands of results per run. No account, no API key.

### Features

- 🔎 **Full-text search** — across titles, authors, journals and more
- 📅 **Date filters** — from/until publication date
- 🗂️ **Type filter** — journal-article, book-chapter, proceedings-article, dataset, preprint…
- 🔗 **DOIs + citation counts** — `is-referenced-by-count` per work, plus reference counts
- 👥 **Authors with ORCID + affiliation**
- 🧾 **Abstracts** — when the publisher deposits them (coverage varies)
- 📊 **Bulk & deep-paginated** — cursor paging scales to hundreds of thousands of records

### What you get

One record per work:

```json
{
  "doi": "10.1089/crispr.2018.29011.rba",
  "title": "Cultivating CRISPR",
  "type": "journal-article",
  "authors": [{ "name": "Rodolphe Barrangou", "given": "Rodolphe", "family": "Barrangou", "orcid": null, "affiliation": [] }],
  "journal": "The CRISPR Journal",
  "publisher": "SAGE Publications",
  "publishedYear": 2018,
  "citationCount": 12,
  "referenceCount": 0,
  "issn": "2573-1599",
  "subjects": [],
  "url": "https://doi.org/10.1089/crispr.2018.29011.rba",
  "abstract": null
}
```

### Input

```json
{
  "query": "crispr gene editing",
  "fromDate": "2020-01-01",
  "type": "journal-article",
  "maxResults": 5000
}
```

| Field | Description |
|-------|-------------|
| `query` | Free-text search (blank = browse by filters only). |
| `fromDate` / `toDate` | Publication date range. |
| `type` | Crossref work type filter. |
| `hasAbstractOnly` | Only works that include an abstract. |
| `maxResults` | Max works to return. |
| `mailto` | Your email — joins Crossref's faster "polite pool". |

### Use cases

- **Literature reviews & bibliometrics** — pull every work on a topic with citation counts
- **Research databases** — build DOI/metadata datasets at scale
- **Citation analysis** — rank works and authors by impact
- **Publisher / journal monitoring** — track new output by type and date

### Pricing

Pay-per-event: a small amount per work scraped. See the **Pricing** tab for current rates.

### FAQ

**Do I need an API key?**
No. This Actor uses the public [Crossref](https://www.crossref.org) REST API — no account, login or API key needed. You can optionally add your email via `mailto` to join Crossref's faster "polite pool".

**How many works can I scrape per run?**
Set `maxResults` as high as you need — cursor paging scales to hundreds of thousands of records per run, so you can pull an entire topic or date range in one go.

**Is scraping Crossref legal?**
Yes. Crossref openly publishes DOI and citation metadata through its REST API, and this Actor reads only those public records and returns them as-is.

**What format is the output?**
Structured JSON — one record per work — exportable as JSON, CSV or Excel. Each record includes DOI, title, authors with ORCID, journal, publisher, publication year, citation count, references and (where deposited) an abstract.

**Can I filter by date or work type?**
Yes. Use `fromDate`/`toDate` for a publication date range, `type` to keep only `journal-article`, `book-chapter`, `dataset`, `preprint` etc., and `hasAbstractOnly` to return only works that include an abstract.

### Related Actors

Building a citation or metadata dataset? These other hipersoft scrapers pair well with this one:

- [OpenAlex Scraper](https://apify.com/hipersoft/openalex-scraper) — 250M+ scholarly works with citations, authors and abstracts
- [Semantic Scholar Scraper](https://apify.com/hipersoft/semantic-scholar-scraper) — papers with citation and influential-citation metrics
- [PubMed Scraper](https://apify.com/hipersoft/pubmed-scraper) — biomedical papers, abstracts and MeSH terms
- [arXiv Papers Scraper](https://apify.com/hipersoft/arxiv-scraper) — preprints with full abstracts and PDF links

### Notes

Uses the public Crossref REST API. Returns Crossref's openly-available metadata as-is. This is an independent tool and is not affiliated with or endorsed by Crossref.

# Actor input Schema

## `query` (type: `string`):

Free-text search across titles, authors, journals and more. Leave blank to browse purely by the filters below.

## `fromDate` (type: `string`):

Earliest publication date (YYYY-MM-DD or YYYY). Optional.

## `toDate` (type: `string`):

Latest publication date (YYYY-MM-DD or YYYY). Optional.

## `type` (type: `string`):

Filter by type, e.g. journal-article, book-chapter, proceedings-article, dataset, posted-content. Optional.

## `hasAbstractOnly` (type: `boolean`):

Return only records that include an abstract.

## `includeAbstract` (type: `boolean`):

Include abstract text when Crossref has it (coverage varies by publisher).

## `maxResults` (type: `integer`):

Maximum works to return (a broad query can match hundreds of thousands).

## `mailto` (type: `string`):

Your email — joins Crossref's faster 'polite pool'. Optional.

## Actor input object example

```json
{
  "query": "crispr gene editing",
  "hasAbstractOnly": false,
  "includeAbstract": true,
  "maxResults": 1000
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "crispr gene editing"
};

// Run the Actor and wait for it to finish
const run = await client.actor("hipersoft/crossref-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "crispr gene editing" }

# Run the Actor and wait for it to finish
run = client.actor("hipersoft/crossref-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "crispr gene editing"
}' |
apify call hipersoft/crossref-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=hipersoft/crossref-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/NbJ4khi3gHhMK8Rmx/builds/jj4LxerU3gsmYOeeX/openapi.json
