Academic Papers Scraper — OpenAlex, 250M+ Works, No Login avatar

Academic Papers Scraper — OpenAlex, 250M+ Works, No Login

Pricing

from $0.50 / 1,000 papers

Go to Apify Store
Academic Papers Scraper — OpenAlex, 250M+ Works, No Login

Academic Papers Scraper — OpenAlex, 250M+ Works, No Login

Search 250M+ scholarly works from OpenAlex — by keyword, author, institution or journal; filter by year, citations, open access and work type. No login, no API key, no captcha. Normalized JSON: DOI, title, authors, venue, year, citations, OA PDF link, abstract.

Pricing

from $0.50 / 1,000 papers

Rating

0.0

(0)

Developer

Petro Pankov

Petro Pankov

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Academic & Research Papers Scraper — OpenAlex (250M+ works, no login)

Search the world's scholarly literature — research papers, journal articles, preprints, book chapters, datasets and reviews — and get clean, structured rows back. Built directly on OpenAlex, an open catalogue of over 250 million works.

No login. No API key. No captcha. No proxies needed.

Why this instead of a Google Scholar scraper

Google Scholar has no public API. Every scraper built on it is reverse-engineering an HTML surface that actively fights automation — which is why those Actors break, rate-limit, and need proxy budgets. OpenAlex publishes the same literature through a documented, open, versioned API with a stable schema. Nothing to fight, nothing to break.

That means: predictable runs, no proxy costs, no captcha failures, and results you can actually schedule against.

What you can do

  • Keyword search across title, abstract and fulltext — "large language models"
  • By author"Yoshua Bengio" (name match, no author IDs needed)
  • By institution"Massachusetts Institute of Technology", "Max Planck"
  • By journal / conference"Nature", "NeurIPS"
  • Look up by DOI & pull citation counts — every result carries a bare doi and a citedByCount, so you can resolve DOIs to metadata or rank papers by citations
  • Filter by year range, minimum citations, open-access-only, and work type
  • Sort by relevance, citation count, or newest first

Output

Each row is normalized:

fielddescription
id, openAlexUrlOpenAlex work ID and URL
doiDOI (bare, no https://doi.org/ prefix)
titleWork title
authors, firstAuthor, authorCountFull author list and convenience fields
venue, publisher, issnJournal / conference and publisher
year, publishedAtPublication year and ISO date
type, languageWork type (article, preprint, …) and language
citedByCount, referencedWorksCount, fwciCitation metrics
isOpenAccess, oaStatus, pdfUrlOpen-access status and direct free-PDF link
landingPageUrlPublisher landing page
concepts, institutions, countriesTopics and affiliations
abstractFull abstract text (when Include abstract is on)

Example

{
"search": "retrieval augmented generation",
"fromYear": 2023,
"minCitations": 10,
"openAccessOnly": true,
"sortBy": "citations",
"maxResults": 200,
"includeAbstract": true
}

Typical uses

  • Literature reviews and systematic reviews
  • Citation analysis and citation-count tracking across a field or author
  • DOI enrichment — resolve a list of DOIs into full titles, authors, venues and metrics
  • Tracking a lab's, author's or institution's output
  • Building datasets of open-access PDFs for downstream NLP
  • Competitive/landscape research in a scientific field

FAQ

Is this a Google Scholar alternative? Yes — it covers the same research-paper literature (and more) through OpenAlex's open API, without Google Scholar's captchas, rate limits or proxy costs. See the comparison above.

Can I look up papers by DOI or get citation counts? Every row carries the doi, citedByCount and referencedWorksCount fields, so you can pull citations, references and DOIs for any matched work.

Does it export research papers as a dataset? Yes — results come back as a flat dataset (JSON, CSV, Excel) ready for literature reviews, citation analysis, or building NLP/RAG corpora.

Can I use it for bibliometrics? Yes — each row carries citation counts, referenced-works counts and OpenAlex's field-weighted citation impact (fwci), so you can run bibliometric and research-impact analysis without a separate metrics source.

Does it cover biomedical / PubMed papers? OpenAlex indexes every discipline, including biomedical and PubMed-indexed literature, so medical and life-science works show up in the same search.

Can I get only open access papers? Yes — set openAccessOnly and every row comes back with isOpenAccess, the oaStatus (gold, green, hybrid, bronze) and a direct pdfUrl where a free full text exists, so you can build a clean list of open-access papers with downloadable PDFs.

Can I scrape preprints? Yes — preprints are a work type in OpenAlex, so you can pull them with the rest of the literature or restrict a run to preprints via the types filter (e.g. arXiv, bioRxiv and other preprint servers indexed by OpenAlex).

Can I use it for a literature review? Yes — it searches the world's scientific literature by keyword, author, journal and year, so you can assemble the full set of relevant papers for a literature review in one run and export them with abstracts, DOIs and citation counts attached.

Which databases does it replace? It taps the same corpus that powers tools built on Crossref, Microsoft Academic Graph, PubMed and Semantic Scholar, unified into one open source.

Notes

  • Results are capped at today's date unless you set To year — OpenAlex contains a small number of records with future publication dates, and this keeps "newest first" meaningful.
  • Requests go through OpenAlex's polite pool (a contact address is sent with each call), which is the usage pattern OpenAlex asks for.
  • Abstracts are stored by OpenAlex as an inverted index for licensing reasons; this Actor reconstructs them into plain text for you.