Google Scholar Scraper
Pricing
from $1.99 / 1,000 google scholar scraper results
Google Scholar Scraper
Scrapes Google Scholar for academic papers. Extracts the full canonical scholar-vertical schema: title, authors (with profile links), publication venue, publisher, year, citation counts, DOI, PDF/HTML/bibtex/abstract URLs, versions count, cluster ID, subjects/keywords, free PDF flag, and more.
Pricing
from $1.99 / 1,000 google scholar scraper results
Rating
0.0
(0)
Developer
Search API
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
9 days ago
Last modified
Categories
Share
Scrapes Google Scholar for academic papers. Extracts the full canonical scholar-vertical schema: title, authors (with profile links), publication venue, publisher, year, citation counts, DOI, PDF/HTML/bibtex/abstract URLs, versions count, cluster ID, subjects/keywords, free PDF flag, and more.
What this Actor collects
The Actor converts Google Scholar results into one clean JSON record per scholarly work, including authors and profile links, venue and publisher, year, citation and version counts, DOI, document links, subjects, keywords, access status, and search provenance when available.
- Uses the input limits and filters below to control the crawl.
- Stores source-backed fields defined by the 48-field dataset schema.
- Omits optional fields when the source does not expose a value instead of writing nulls or fabricated placeholders.
Use cases
- Literature and catalog research
- Entity and citation enrichment
- Specialized-source monitoring
Input
Provide input in JSON. Fields marked required must be supplied. The Default / example column shows a schema default when one exists; otherwise it shows a documented prefill or fixture value.
| Field | Type | Required | Default / example | Description |
|---|---|---|---|---|
query | string | Yes | "machine learning" | The academic search query to look up on Google Scholar |
maxItems | integer | No | 50 | Maximum number of scholar results to retrieve |
yearFrom | integer | No | — | Filter results published from this year onwards (optional) |
yearTo | integer | No | — | Filter results published up to this year (optional) |
sortBy | string | No | "relevance" | Sort results by relevance or date |
proxyConfiguration | object | No | {"useApifyProxy":true,"apifyProxyGroups":["RESIDENTIAL"]} | Proxy settings for the scraper. Use Apify residential proxies to avoid blocks. |
Example input
{"query": "transformer neural network","maxItems": 50,"proxyConfiguration": {"useApifyProxy": true,"apifyProxyGroups": ["GOOGLE_SERP"]},"sortBy": "relevance"}
Output
The default dataset contains one item per scholarly publication result. The following are the most useful fields; DOI, citation links, versions, document URLs, subjects, and access fields depend on the Scholar result.
| Field | Type | Description |
|---|---|---|
position | integer | Position |
title | string | Title |
authors | array | Authors |
publication | string | Publication |
year | integer | Year |
citedBy | integer | Cited By |
searchQuery | string | Search Query |
scrapedAt | string | Scraped At |
type | string | Type |
description | string | Description |
snippet | string | Snippet |
url | string | URL |
link | string | Link |
query | string | Query |
resultType | string | Result Type |
page | integer | Page |
Example dataset item
This compact example is taken from local Actor storage. Long text and nested collections are shortened for documentation only.
{"position": 1,"title": "Gradient-based learning applied to document recognition","authors": ["Yann LeCun","Léon Bottou"],"publication": "Proceedings of the IEEE","year": 1998,"citedBy": 58655,"searchQuery": "transformer neural network","scrapedAt": "2026-07-23T11:51:14.711Z","type": "scholar","description": "Multilayer neural networks trained with the back-propagation algorithm constitute the best example of a successful gradient based learning technique. Given an appropriate network …","snippet": "Multilayer neural networks trained with the back-propagation algorithm constitute the best example of a successful gradient based learning technique. Given an appropriate network …","url": "https://doi.org/10.1109/5.726791"}
