Academic Papers Scraper — OpenAlex, 250M+ Works, No Login
Pricing
from $0.50 / 1,000 papers
Academic Papers Scraper — OpenAlex, 250M+ Works, No Login
Search 250M+ scholarly works from OpenAlex — by keyword, author, institution or journal; filter by year, citations, open access and work type. No login, no API key, no captcha. Normalized JSON: DOI, title, authors, venue, year, citations, OA PDF link, abstract.
Pricing
from $0.50 / 1,000 papers
Rating
0.0
(0)
Developer
Petro Pankov
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Share
Academic & Research Papers Scraper — OpenAlex (250M+ works, no login)
Search the world's scholarly literature — research papers, journal articles, preprints, book chapters, datasets and reviews — and get clean, structured rows back. Built directly on OpenAlex, an open catalogue of over 250 million works.
No login. No API key. No captcha. No proxies needed.
Why this instead of a Google Scholar scraper
Google Scholar has no public API. Every scraper built on it is reverse-engineering an HTML surface that actively fights automation — which is why those Actors break, rate-limit, and need proxy budgets. OpenAlex publishes the same literature through a documented, open, versioned API with a stable schema. Nothing to fight, nothing to break.
That means: predictable runs, no proxy costs, no captcha failures, and results you can actually schedule against.
What you can do
- Keyword search across title, abstract and fulltext —
"large language models" - By author —
"Yoshua Bengio"(name match, no author IDs needed) - By institution —
"Massachusetts Institute of Technology","Max Planck" - By journal / conference —
"Nature","NeurIPS" - Look up by DOI & pull citation counts — every result carries a bare
doiand acitedByCount, so you can resolve DOIs to metadata or rank papers by citations - Filter by year range, minimum citations, open-access-only, and work type
- Sort by relevance, citation count, or newest first
Output
Each row is normalized:
| field | description |
|---|---|
id, openAlexUrl | OpenAlex work ID and URL |
doi | DOI (bare, no https://doi.org/ prefix) |
title | Work title |
authors, firstAuthor, authorCount | Full author list and convenience fields |
venue, publisher, issn | Journal / conference and publisher |
year, publishedAt | Publication year and ISO date |
type, language | Work type (article, preprint, …) and language |
citedByCount, referencedWorksCount, fwci | Citation metrics |
isOpenAccess, oaStatus, pdfUrl | Open-access status and direct free-PDF link |
landingPageUrl | Publisher landing page |
concepts, institutions, countries | Topics and affiliations |
abstract | Full abstract text (when Include abstract is on) |
Example
{"search": "retrieval augmented generation","fromYear": 2023,"minCitations": 10,"openAccessOnly": true,"sortBy": "citations","maxResults": 200,"includeAbstract": true}
Typical uses
- Literature reviews and systematic reviews
- Citation analysis and citation-count tracking across a field or author
- DOI enrichment — resolve a list of DOIs into full titles, authors, venues and metrics
- Tracking a lab's, author's or institution's output
- Building datasets of open-access PDFs for downstream NLP
- Competitive/landscape research in a scientific field
FAQ
Is this a Google Scholar alternative? Yes — it covers the same research-paper literature (and more) through OpenAlex's open API, without Google Scholar's captchas, rate limits or proxy costs. See the comparison above.
Can I look up papers by DOI or get citation counts? Every row carries the doi, citedByCount
and referencedWorksCount fields, so you can pull citations, references and DOIs for any matched work.
Does it export research papers as a dataset? Yes — results come back as a flat dataset (JSON, CSV, Excel) ready for literature reviews, citation analysis, or building NLP/RAG corpora.
Can I use it for bibliometrics? Yes — each row carries citation counts, referenced-works counts
and OpenAlex's field-weighted citation impact (fwci), so you can run bibliometric and
research-impact analysis without a separate metrics source.
Does it cover biomedical / PubMed papers? OpenAlex indexes every discipline, including biomedical and PubMed-indexed literature, so medical and life-science works show up in the same search.
Can I get only open access papers? Yes — set openAccessOnly and every row comes back with
isOpenAccess, the oaStatus (gold, green, hybrid, bronze) and a direct pdfUrl where a free
full text exists, so you can build a clean list of open-access papers with downloadable PDFs.
Can I scrape preprints? Yes — preprints are a work type in OpenAlex, so you can pull them
with the rest of the literature or restrict a run to preprints via the types filter (e.g. arXiv,
bioRxiv and other preprint servers indexed by OpenAlex).
Can I use it for a literature review? Yes — it searches the world's scientific literature by keyword, author, journal and year, so you can assemble the full set of relevant papers for a literature review in one run and export them with abstracts, DOIs and citation counts attached.
Which databases does it replace? It taps the same corpus that powers tools built on Crossref, Microsoft Academic Graph, PubMed and Semantic Scholar, unified into one open source.
Notes
- Results are capped at today's date unless you set To year — OpenAlex contains a small number of records with future publication dates, and this keeps "newest first" meaningful.
- Requests go through OpenAlex's polite pool (a contact address is sent with each call), which is the usage pattern OpenAlex asks for.
- Abstracts are stored by OpenAlex as an inverted index for licensing reasons; this Actor reconstructs them into plain text for you.