Google Scholar Scraper avatar

Google Scholar Scraper

Pricing

from $1.99 / 1,000 google scholar scraper results

Go to Apify Store
Google Scholar Scraper

Google Scholar Scraper

Scrapes Google Scholar for academic papers. Extracts the full canonical scholar-vertical schema: title, authors (with profile links), publication venue, publisher, year, citation counts, DOI, PDF/HTML/bibtex/abstract URLs, versions count, cluster ID, subjects/keywords, free PDF flag, and more.

Pricing

from $1.99 / 1,000 google scholar scraper results

Rating

0.0

(0)

Developer

Search API

Search API

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 days ago

Last modified

Share

Scrapes Google Scholar for academic papers. Extracts the full canonical scholar-vertical schema: title, authors (with profile links), publication venue, publisher, year, citation counts, DOI, PDF/HTML/bibtex/abstract URLs, versions count, cluster ID, subjects/keywords, free PDF flag, and more.

What this Actor collects

The Actor converts Google Scholar results into one clean JSON record per scholarly work, including authors and profile links, venue and publisher, year, citation and version counts, DOI, document links, subjects, keywords, access status, and search provenance when available.

  • Uses the input limits and filters below to control the crawl.
  • Stores source-backed fields defined by the 48-field dataset schema.
  • Omits optional fields when the source does not expose a value instead of writing nulls or fabricated placeholders.

Use cases

  • Literature and catalog research
  • Entity and citation enrichment
  • Specialized-source monitoring

Input

Provide input in JSON. Fields marked required must be supplied. The Default / example column shows a schema default when one exists; otherwise it shows a documented prefill or fixture value.

FieldTypeRequiredDefault / exampleDescription
querystringYes"machine learning"The academic search query to look up on Google Scholar
maxItemsintegerNo50Maximum number of scholar results to retrieve
yearFromintegerNoFilter results published from this year onwards (optional)
yearTointegerNoFilter results published up to this year (optional)
sortBystringNo"relevance"Sort results by relevance or date
proxyConfigurationobjectNo{"useApifyProxy":true,"apifyProxyGroups":["RESIDENTIAL"]}Proxy settings for the scraper. Use Apify residential proxies to avoid blocks.

Example input

{
"query": "transformer neural network",
"maxItems": 50,
"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyGroups": [
"GOOGLE_SERP"
]
},
"sortBy": "relevance"
}

Output

The default dataset contains one item per scholarly publication result. The following are the most useful fields; DOI, citation links, versions, document URLs, subjects, and access fields depend on the Scholar result.

FieldTypeDescription
positionintegerPosition
titlestringTitle
authorsarrayAuthors
publicationstringPublication
yearintegerYear
citedByintegerCited By
searchQuerystringSearch Query
scrapedAtstringScraped At
typestringType
descriptionstringDescription
snippetstringSnippet
urlstringURL
linkstringLink
querystringQuery
resultTypestringResult Type
pageintegerPage

Example dataset item

This compact example is taken from local Actor storage. Long text and nested collections are shortened for documentation only.

{
"position": 1,
"title": "Gradient-based learning applied to document recognition",
"authors": [
"Yann LeCun",
"Léon Bottou"
],
"publication": "Proceedings of the IEEE",
"year": 1998,
"citedBy": 58655,
"searchQuery": "transformer neural network",
"scrapedAt": "2026-07-23T11:51:14.711Z",
"type": "scholar",
"description": "Multilayer neural networks trained with the back-propagation algorithm constitute the best example of a successful gradient based learning technique. Given an appropriate network …",
"snippet": "Multilayer neural networks trained with the back-propagation algorithm constitute the best example of a successful gradient based learning technique. Given an appropriate network …",
"url": "https://doi.org/10.1109/5.726791"
}