# Shopify Store Products Scraper (`centspy/shopify-products-scraper`) Actor

Scrape any Shopify store's full product catalog via the public products.json feed. No proxy, no browser.

- **URL**: https://apify.com/centspy/shopify-products-scraper.md
- **Developed by:** [brandon nadeau](https://apify.com/centspy) (community)
- **Categories:** E-commerce, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 product scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Shopify Store Products Scraper

Extract the **full product catalog** from any Shopify store in seconds — titles, prices, variants, SKUs, inventory status, vendors, tags, and image URLs. Perfect for competitor research, dropshipping product discovery, price monitoring, and catalog analysis.

This scraper reads Shopify's **public `products.json` feed**, which every Shopify storefront exposes by default. That means:

- **No proxies needed** — keeps your cost low.
- **No browser** — fast and lightweight.
- **No broken selectors** — it reads a stable JSON feed, not fragile HTML, so it doesn't break when stores change their design.

### What it does

Give it one or more Shopify store URLs and it returns every product, fully structured and ready to drop into a spreadsheet, database, or your own app.

### Input

| Field | Type | Description |
|-------|------|-------------|
| `storeUrls` | array | One or more Shopify store URLs. The domain alone is enough (e.g. `allbirds.com`). **Required.** |
| `collectionHandle` | string | Optional. Scope the scrape to a single collection (e.g. `mens-shoes`). Leave blank for the full catalog. |
| `maxProductsPerStore` | integer | Cap products per store. `0` = no limit. Default `0`. |
| `includeVariants` | boolean | Include full variant details (size/color/SKU/price). Default `true`. |
| `includeImages` | boolean | Include image URLs. Default `true`. |

#### Example input

```json
{
    "storeUrls": ["https://www.allbirds.com", "https://shop.gymshark.com"],
    "maxProductsPerStore": 500,
    "includeVariants": true,
    "includeImages": true
}
```

### Output

Each product is returned as a structured row:

```json
{
    "store": "https://www.allbirds.com",
    "id": 6789012345,
    "title": "Men's Wool Runners",
    "handle": "mens-wool-runners",
    "url": "https://www.allbirds.com/products/mens-wool-runners",
    "vendor": "Allbirds",
    "productType": "Shoes",
    "tags": ["wool", "running"],
    "priceMin": 98.0,
    "priceMax": 110.0,
    "available": true,
    "variantCount": 12,
    "variants": [
        {
            "id": 39812345678,
            "title": "Natural Grey / 9",
            "sku": "WR-NG-9",
            "price": 98.0,
            "compareAtPrice": null,
            "available": true,
            "option1": "Natural Grey",
            "option2": "9",
            "option3": null
        }
    ],
    "images": ["https://cdn.shopify.com/..."],
    "featuredImage": "https://cdn.shopify.com/..."
}
```

Export to **CSV, JSON, Excel, or HTML** from the dataset, or pull it via the Apify API.

### Pricing

This Actor uses **pay-per-event** pricing — you pay a small amount per product scraped. Platform usage is included in the price, so what you see is what you pay. A typical run scraping 1,000 products costs roughly the price of one event × 1,000.

### Notes & limitations

- A small number of stores disable their public `products.json` feed. For those, the Actor logs a warning and moves on.
- The feed returns published, customer-facing products. It does not include draft or hidden products.
- Inventory `available` reflects each variant's purchasable status at scrape time.

### Use cases

- **Competitor monitoring** — track a rival's catalog, pricing, and new launches.
- **Dropshipping** — discover products and suppliers across stores.
- **Price intelligence** — feed pricing into your own monitoring dashboards.
- **Market research** — analyze assortment, vendors, and product types at scale.

# Actor input Schema

## `storeUrls` (type: `array`):

One or more Shopify store URLs. The domain is enough (e.g. allbirds.com) — paths are ignored.

## `collectionHandle` (type: `string`):

Scope the scrape to a single collection across all stores, e.g. 'mens-shoes'. Leave blank to scrape the entire catalog.

## `maxProductsPerStore` (type: `integer`):

Cap the number of products scraped per store. Set 0 for no limit (scrape the whole catalog).

## `includeVariants` (type: `boolean`):

Include the full variant list (size/color/SKU/price) for each product.

## `includeImages` (type: `boolean`):

Include image URLs for each product.

## Actor input object example

```json
{
  "storeUrls": [
    "https://www.allbirds.com",
    "https://shop.gymshark.com"
  ],
  "collectionHandle": "",
  "maxProductsPerStore": 0,
  "includeVariants": true,
  "includeImages": true
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "storeUrls": [
        "https://www.allbirds.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("centspy/shopify-products-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "storeUrls": ["https://www.allbirds.com"] }

# Run the Actor and wait for it to finish
run = client.actor("centspy/shopify-products-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "storeUrls": [
    "https://www.allbirds.com"
  ]
}' |
apify call centspy/shopify-products-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=centspy/shopify-products-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/0ce0MWRipjKaPWdPI/builds/ZYY9EnEWSFIwJ5UiP/openapi.json
