Shopify Products Scraper
Pricing
from $0.50 / 1,000 products
Shopify Products Scraper
Extract products, prices, variants and stock from any public Shopify store via its /products.json endpoint. Flat output, no browser. Unofficial, not affiliated with Shopify.
Pricing
from $0.50 / 1,000 products
Rating
0.0
(0)
Developer
Bluefin
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Shopify Products Scraper
This is an unofficial tool. It is not affiliated with, endorsed by, or sponsored by Shopify Inc. "Shopify" is a trademark of Shopify Inc.
Extract the full product catalog from any public Shopify store — every product and, if you want, every variant — as clean, flat rows you can download as JSON, CSV, or Excel. It reads the store's public /products.json endpoint, so it is fast and does not run a browser.
Point it at one or more store URLs (for example https://www.allbirds.com) and it walks the whole catalog page by page, then hands you titles, handles, vendors, product types, tags, prices, compare-at prices, SKUs, availability, images, and canonical product URLs. Running it on the Apify platform adds scheduled runs, API access, proxy rotation, and integrations with Make, Zapier, Google Sheets, and more.
Why use this scraper?
- Competitor and price monitoring — track a competitor's catalog, price points, and what they mark as on sale (via
compare_at_price). - Catalog research — pull an entire store's assortment for analysis, dropshipping research, or building a product feed.
- Availability tracking — see which variants are in stock at scrape time.
- Lightweight and cheap — HTTP-only (no headless browser), so runs are quick and use little compute.
How to use it
- Add one or more Shopify store URLs to Start URLs. You can paste the store home page (
https://store.com) or the fullhttps://store.com/products.jsonpath — both work. - Optionally set Max products per store, toggle Expand variants into separate rows, or enable Also fetch collections.
- Leave Proxy configuration on Apify Proxy (recommended — see the note under Limitations).
- Click Start. When the run finishes, open the Output tab or export the dataset as JSON, CSV, or Excel.
Input
| Field | Type | Default | Description |
|---|---|---|---|
startUrls | array | (required) | Shopify store URLs. Home page or /products.json path both accepted. |
maxProductsPerStore | integer | 0 | Cap on products per store. 0 means no limit. Counted per product, not per variant. |
includeVariants | boolean | true | Expand every variant into its own row. When false, one row per product using the default variant. |
includeCollections | boolean | false | Also fetch the store's public collections list into a separate collections dataset. |
proxyConfiguration | object | Apify Proxy | Proxy settings. |
Example input:
{"startUrls": [{ "url": "https://www.allbirds.com" }],"maxProductsPerStore": 0,"includeVariants": true,"includeCollections": false,"proxyConfiguration": { "useApifyProxy": true }}
Output
Each row is one product variant (or one product when includeVariants is false). Real example row:
{"store_url": "https://greatjonesgoods.com","product_id": 8410904559695,"handle": "big-deal-saucy","title": "Big Deal & Saucy","vendor": "Great Jones","product_type": "Bundle","tags": ["badge_Save $60!"],"created_at": "2026-03-16T16:05:51-04:00","updated_at": "2026-08-01T19:16:49-04:00","published_at": "2026-03-30T14:47:52-04:00","product_url": "https://greatjonesgoods.com/products/big-deal-saucy","price": "190.00","compare_at_price": "250.00","variant_id": 44838392430671,"variant_title": "Default Title","sku": null,"available": true,"image_url": "https://cdn.shopify.com/s/files/1/0066/9312/6202/files/BigDeal_Saucy.png"}
You can download the dataset in JSON, HTML, CSV, or Excel from the Output tab or the API.
Data fields
| Field | Description |
|---|---|
store_url | Origin of the store the row came from. |
product_id | Shopify product ID. |
handle | URL slug of the product. |
title | Product title. |
vendor | Vendor / brand as set in Shopify. |
product_type | Shopify product type. |
tags | Array of product tags. |
created_at / updated_at / published_at | Product timestamps. |
price | Variant price (string, as Shopify returns it). |
compare_at_price | Compare-at (list) price, or null. |
variant_id | Shopify variant ID (null if the product has no variants). |
variant_title | Variant name, e.g. "Small" or "Default Title". |
sku | Variant SKU, or null if the store did not set one. |
available | Whether the variant was in stock at scrape time. |
image_url | Variant image if present, otherwise the first product image. |
product_url | Canonical product page URL. |
With includeCollections enabled, the store's collections are written to a named dataset called collections, which you can retrieve from the API or the run's Storage tab. It holds collection_id, handle, title, description, products_count, published_at, updated_at, image_url, and collection_url. Note that named datasets may not appear in the Console Output tab (which shows the default dataset) — open the run's Storage tab, or the collections dataset via the API, to see them.
Pricing
This Actor uses pay-per-event pricing with two charges:
- Run start — charged once each time a run starts (Apify's built-in start event).
- Product scraped — charged once per product stored. Variants of the same product do not add extra charges, so turning variant expansion on does not increase cost.
A store with 500 products is billed as one run start plus 500 product charges, no matter how many variants those products have. So a product with 20 size/color variants is billed once, not 20 times. Check the Actor's Store page for the current per-charge amounts.
The run-start charge is counted per gigabyte of memory: it is charged once for runs up to and including 1 GB of RAM, then once more for each additional GB (for example, 4 GB of RAM is billed as 4 start charges). This Actor runs HTTP-only and sets a 1 GB default, so under the default it is a single start charge.
For developers / publishers — important, read before publishing. Billing is per product, achieved with a single custom event named
product-result(defined insrc/main.tsand.actor/pay_per_event.json). Apify's pay-per-event model also offers two automatic "synthetic" events, and one of them is a trap for this Actor:
- Use the synthetic
apify-actor-startfor run startup (Apify charges it automatically — do not charge it from code, that fails).- You must remove the synthetic
apify-default-dataset-itemevent in the Monetization wizard. If left enabled, it auto-charges once per row in the default dataset — and this Actor writes one row per variant, which would bill a 20-variant product 20 times and make the "no extra charge for variants" promise above false.Pay-per-event pricing is configured in the Apify Console Monetization wizard, not in
actor.json. FollowPUBLISH_CHECKLIST.mdexactly, and after publishing run the billing verification step there to confirm no per-row charges appear.
Limitations (please read)
- Public data only. It reads the store's public
/products.json, the same JSON any browser can request. It cannot access password-protected stores, unpublished products, admin-only fields, or exact inventory counts. - Password-protected, closed, or non-Shopify sites are skipped, not scraped. Each skip is logged with a reason and recorded in the
SUMMARYrecord in the run's key-value store. If every store is unreachable, the run fails rather than reporting a false success. - Some stores block datacenter IPs. Many Shopify stores sit behind Cloudflare and rate-limit or block datacenter traffic (HTTP 429/403). Keeping Apify Proxy enabled (the default) rotates IPs and materially improves success rates. If a store still returns 429, try residential proxy groups.
- No currency field. The
/products.jsonendpoint does not include a currency code, so prices are returned as-is without one. - No product-to-collection mapping. With
includeCollections, you get collection metadata (title, handle, product count), not which products belong to which collection.
FAQ
Is scraping Shopify stores legal?
This Actor only requests the store's own public /products.json endpoint — publicly available data. You are responsible for using the output in line with the target store's terms and applicable laws. It targets product catalog data only and does not collect customer or account data. Note that catalog fields such as vendor can occasionally contain a person's name (for example, a sole trader's brand), so treat the output accordingly.
Why did a store return 0 products or get skipped?
The store may be password-protected, closed, not a Shopify store, or blocking your IP. Check the SUMMARY record in the run's key-value store for the exact per-store reason, and try enabling Apify Proxy.
Does it get every product?
It paginates until the store returns an empty page, so it retrieves the full public catalog unless you set maxProductsPerStore.
How do I report a problem? Use the Issues tab on the Actor page.