# Europe PMC Scraper — Life-Science Papers & Full Text (`hipersoft/europepmc-scraper`) Actor

Bulk-scrape life-science and biomedical literature from Europe PMC (40M+ records: PubMed, PMC, preprints, patents): title, abstract, authors, journal, year, DOI, PMID/PMCID, citation count, MeSH terms, open-access flag and full-text/PDF links. Thousands per run. No login, no API key.

- **URL**: https://apify.com/hipersoft/europepmc-scraper.md
- **Developed by:** [hiper soft](https://apify.com/hipersoft) (community)
- **Categories:** Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.0004 / article scraped

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Europe PMC Scraper — Life-Science Papers & Full Text

Bulk-scrape biomedical and life-science literature from **Europe PMC** — a 40M+ record index spanning **PubMed, PubMed Central, preprints, patents and agricultural literature**. Get titles, abstracts, authors, journals, years, DOIs, PMIDs/PMCIDs, **citation counts**, MeSH terms, open-access flags and **direct full-text / PDF links**. Thousands of records per run. No account, no API key.

### Features

- 🔎 **Broad life-science coverage** — PubMed + PMC + preprints + patents in one search
- 📄 **Full-text & PDF links** — direct links where the article is openly available
- 📈 **Citation counts** — `citedByCount` per article
- 🏷️ **Rich metadata** — MeSH terms, keywords, publication types, journal, IDs (DOI/PMID/PMCID)
- 🔓 **Open-access filter** — restrict to freely readable articles
- 📊 **Bulk & cursor-paginated** — scales to hundreds of thousands of records per query

### What you get

One record per article:

```json
{
  "id": "38512345",
  "source": "MED",
  "pmid": "38512345",
  "pmcId": "PMC10987654",
  "doi": "10.1038/s41586-024-00000-0",
  "title": "…",
  "authors": ["Jane A Smith", "…"],
  "journal": "Nature",
  "pubYear": 2024,
  "citedByCount": 37,
  "isOpenAccess": true,
  "hasPdf": true,
  "pdfUrl": "https://europepmc.org/articles/PMC10987654?pdf=render",
  "fullTextUrl": "https://europepmc.org/articles/PMC10987654",
  "meshTerms": ["…"],
  "url": "https://europepmc.org/article/MED/38512345",
  "abstract": "…"
}
```

### Input

```json
{
  "query": "crispr AND cancer",
  "openAccessOnly": false,
  "maxResults": 5000
}
```

| Field | Description |
|-------|-------------|
| `query` | Europe PMC query (supports fields like `AUTH:`, `JOURNAL:`, `PUB_YEAR:`, booleans). |
| `openAccessOnly` | Only open-access articles. |
| `includeAbstract` | Include abstract text. |
| `maxResults` | Max articles to return. |

### Use cases

- **Systematic & literature reviews** — every relevant paper with abstracts and citations
- **Full-text mining** — grab open-access PDF/HTML links at scale
- **Bibliometrics** — rank by citations across PubMed + PMC + preprints
- **Biomedical datasets** — build corpora with MeSH-indexed metadata

### Pricing

Pay-per-event: a small amount per article scraped. See the **Pricing** tab for current rates.

### FAQ

**Do I need an API key?**
No. This Actor uses the public [Europe PMC](https://europepmc.org) REST API — no account, login or API key required.

**How many articles can I scrape per run?**
Set `maxResults` as high as you need — the API is cursor-paginated and scales to hundreds of thousands of records per query.

**Is scraping Europe PMC legal?**
Yes. Europe PMC openly publishes its life-science index through a public REST API, and this Actor reads only those openly available records and returns them as-is.

**What format is the output?**
Structured JSON — one record per article — exportable as JSON, CSV or Excel. Each record includes title, abstract, authors, journal, year, DOI, PMID/PMCID, citation count, MeSH terms, open-access flags and direct full-text/PDF links.

**Can I filter by open access or use field search?**
Yes. Flip `openAccessOnly` to keep only freely readable articles, and the `query` field supports Europe PMC syntax like `AUTH:`, `JOURNAL:`, `PUB_YEAR:` and booleans.

### Related Actors

Building a biomedical literature dataset? These other hipersoft scrapers pair well with this one:

- [PubMed Scraper](https://apify.com/hipersoft/pubmed-scraper) — biomedical papers, abstracts and MeSH terms from the NLM index
- [ClinicalTrials.gov Scraper](https://apify.com/hipersoft/clinicaltrials-scraper) — bulk clinical trial records, sponsors and locations
- [OpenAlex Scraper](https://apify.com/hipersoft/openalex-scraper) — 250M+ scholarly works with citations and abstracts
- [Crossref Scraper](https://apify.com/hipersoft/crossref-scraper) — DOIs, citation counts and metadata from 150M+ works

### Notes

Uses the public Europe PMC REST API. Returns its openly-available metadata as-is. This is an independent tool and is not affiliated with or endorsed by Europe PMC or EMBL-EBI.

# Actor input Schema

## `query` (type: `string`):

Europe PMC search query. Supports fields and booleans, e.g. "malaria AND vaccine", "AUTH:"Doudna"", "cancer AND PUB\_YEAR:2023".

## `openAccessOnly` (type: `boolean`):

Return only open-access articles (full text freely available).

## `includeAbstract` (type: `boolean`):

Include the abstract text for each article.

## `maxResults` (type: `integer`):

Maximum articles to return (a broad query matches hundreds of thousands).

## Actor input object example

```json
{
  "query": "crispr AND cancer",
  "openAccessOnly": false,
  "includeAbstract": true,
  "maxResults": 1000
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "crispr AND cancer"
};

// Run the Actor and wait for it to finish
const run = await client.actor("hipersoft/europepmc-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "crispr AND cancer" }

# Run the Actor and wait for it to finish
run = client.actor("hipersoft/europepmc-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "crispr AND cancer"
}' |
apify call hipersoft/europepmc-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=hipersoft/europepmc-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bimsjcfe3mhVqe13C/builds/hJF64WNdepVPNoack/openapi.json
