# Google Patents Scraper (`masked_hacker/google-patents-scraper`) Actor

Scrape structured patent records (title, abstract, claims, inventors, assignee, dates) from Google Patents.

- **URL**: https://apify.com/masked\_hacker/google-patents-scraper.md
- **Developed by:** [Masked Hacker](https://apify.com/masked_hacker) (community)
- **Categories:** Other
- **Stats:** 2 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 patents

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Google Patents Scraper: Titles, Abstracts, Claims & Assignees

Turn any [Google Patents](https://patents.google.com) search into a clean, structured dataset of
patents. Search by keyword and narrow by inventor, assignee or priority-date range, and get
**title, abstract, claims, inventors, assignee, filing/publication dates and CPC codes** for every
result, fully paginated.

Perfect for **patent landscaping, competitor IP monitoring, prior-art search, and technology
scouting** across any field.

### What you get

- 🔎 **Search by keyword**, narrowed by inventor, assignee, or priority-date range.
- 📄 **Full record per patent:** title, abstract, and every claim.
- 🏢 **Inventors and assignee** (the patent owner), in canonical English.
- 🗓️ **Priority, filing, publication and grant dates**, plus application number.
- 🏷️ **CPC classification codes** and issuing-authority country code.
- 📊 Clean JSON / CSV / Excel export, ready for a spreadsheet or a pipeline.

### Input

| Field | Type | Description |
|---|---|---|
| `searchQueries` | string\[] | Keyword searches to run, e.g. `self-driving vehicle`. Required. |
| `inventor` | string | Restrict every query to this inventor name. |
| `assignee` | string | Restrict every query to this assignee (patent owner), e.g. `Waymo`. |
| `dateFrom` / `dateTo` | string | Priority-date range, `YYYY-MM-DD`. |
| `maxItems` | int | Stop after this many patents across all queries (default 100). |
| `proxyConfiguration` | proxy | Datacenter rotation by default; switch to residential only if rate-limited. |

#### Example input

```json
{
  "searchQueries": ["autonomous vehicle"],
  "assignee": "Waymo",
  "dateFrom": "2020-01-01",
  "maxItems": 100
}
```

### Output

One record per patent, deduplicated by publication number.

| Field | Description |
|---|---|
| `publicationNumber`, `title`, `url` | Patent identity and its Google Patents page. |
| `abstract`, `claims` | Abstract text and one entry per claim. |
| `inventors`, `assignees` | Inventor name(s) and original assignee(s). |
| `priorityDate`, `filingDate`, `publicationDate`, `grantDate` | Key dates. |
| `applicationNumber`, `countryCode`, `cpcCodes` | Application number, authority, and CPC codes. |
| `pdfUrl`, `snippet` | PDF link and the search snippet that matched. |
| `query`, `scrapedAt` | Source query and scrape time. |

#### Example output

```json
{
  "publicationNumber": "US12525077B2",
  "title": "Self-driving vehicles and weigh station operation",
  "url": "https://patents.google.com/patent/US12525077B2/en",
  "abstract": "The technology involves operation of a self-driving truck...",
  "inventors": ["Vijaysai Patnaik"],
  "assignees": ["Waymo LLC"],
  "priorityDate": "2019-12-16",
  "filingDate": "2023-11-14",
  "publicationDate": "2026-01-13",
  "cpcCodes": ["B60W50/0205"],
  "query": "self-driving vehicle"
}
```

### FAQ

**Do I need a Google account or API key?** No. Just provide a search query.

**How many results per query?** Google Patents pages results in tens and caps a single query at
1000 results; the actor paginates up to your `maxItems` cap.

**Why are some abstracts or claims empty?** Patents from some non-US offices don't expose an
English abstract or machine-readable claims on Google Patents; core bibliographic fields are always
populated.

# Actor input Schema

## `searchQueries` (type: `array`):

Keyword searches to run on Google Patents. Each entry is one search, paginated up to the Max patents cap.

## `inventor` (type: `string`):

Optional. Restrict every query to this inventor name.

## `assignee` (type: `string`):

Optional. Restrict every query to this assignee (patent owner), e.g. "Waymo".

## `dateFrom` (type: `string`):

Optional. Only include patents with a priority date on or after this date (YYYY-MM-DD).

## `dateTo` (type: `string`):

Optional. Only include patents with a priority date on or before this date (YYYY-MM-DD).

## `maxItems` (type: `integer`):

Stop after this many patent records across all queries (endpoint caps each query at 1000).

## `proxyConfiguration` (type: `object`):

Datacenter proxy rotation by default is enough — the endpoints only rate-limit. Switch to a residential group only if rate-limiting persists.

## Actor input object example

```json
{
  "searchQueries": [
    "battery management system",
    "mRNA vaccine"
  ],
  "maxItems": 100,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All scraped patent records in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "self-driving vehicle"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("masked_hacker/google-patents-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchQueries": ["self-driving vehicle"] }

# Run the Actor and wait for it to finish
run = client.actor("masked_hacker/google-patents-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "self-driving vehicle"
  ]
}' |
apify call masked_hacker/google-patents-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=masked_hacker/google-patents-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/c79WA4HQop6mkyZ3M/builds/oTg8jW6qfwF3CIsPF/openapi.json
