# Zara Product Scraper — Prices, Stock & Variants (`khadinakbar/zara-product-scraper`) Actor

Extract validated Zara product data from product URLs, search terms, or category pages. Get prices, sale prices, availability, images, color and size options, descriptions, and canonical URLs for fashion research and monitoring. Pay only for complete product rows.

- **URL**: https://apify.com/khadinakbar/zara-product-scraper.md
- **Developed by:** [Khadin Akbar](https://apify.com/khadinakbar) (community)
- **Categories:** E-commerce, Automation, MCP servers
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 product scrapeds

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Zara Product Scraper — Prices, Stock & Variants

Zara Product Scraper is an Apify Actor for fashion analysts, ecommerce teams, monitoring workflows, and AI agents that need validated public Zara product records. It accepts Zara product URLs, category pages, search URLs, or product search terms. Each saved dataset item represents one complete product record, with fields such as productId, reference when available, name, price, originalPrice when shown, currency, availability, colors, sizes, images, description, category, canonical URL, source type, search term, and scrapedAt. The outcome is a clean, one-record-per-product dataset built for downstream analysis and automation.

### Best fit and connected workflows

This Actor fits workflows that start with a public Zara product page, a Zara category page, or a Zara search page and need structured product rows. It also fits discovery runs where a simple term such as `linen shirt` is turned into a Zara search page for collection.

Common routing patterns include:

- Direct product URL monitoring for price, sale price, stock text, and variants.
- Search-term discovery for broader product collection from a selected Zara storefront.
- Category-page collection for a bounded slice of a public Zara assortment.
- Apify MCP usage when an agent needs a focused Zara product extraction tool with validated rows and provenance.

### Practical scenario

Maya, a merchandiser, starts with a Zara search URL for linen shirts and sets `maxProducts` to 20. The run returns rows with `productId`, `name`, `price`, `availability`, `colors`, `sizes`, `images`, `url`, and `scrapedAt`. Maya uses `availability` and `originalPrice` to review live assortment and markdowns, then opens the canonical URLs to compare the storefront pages.

### Input fields

| Field | Type | Purpose |
| --- | --- | --- |
| `startUrls` | array of strings | Zara product, category, or search URLs for direct collection |
| `searchTerms` | array of strings | Product terms used to construct Zara search pages |
| `country` | string | Storefront used when building search URLs from `searchTerms` |
| `maxProducts` | integer | Global output and billing cap for the run |
| `maxScrollsPerListing` | integer | Scroll count for lazy-loaded search and category pages |
| `includeVariants` | boolean | Includes color and size options from the product page |
| `proxyConfiguration` | object | Proxy settings, with Residential as the default group |

#### Focused JSON example

```json
{
  "startUrls": [
    "https://www.zara.com/us/en/search?searchTerm=linen"
  ],
  "searchTerms": [
    "linen shirt"
  ],
  "country": "US",
  "maxProducts": 10,
  "maxScrollsPerListing": 2,
  "includeVariants": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

### Output fields

| Field | Type | Meaning |
| --- | --- | --- |
| `productId` | string | Stable Zara product identifier when exposed |
| `reference` | string or null | Product reference or SKU when shown |
| `name` | string | Current Zara product title |
| `description` | string or null | Product description from the public product page |
| `price` | number or null | Current public price |
| `originalPrice` | number or null | Previous public price when a markdown is shown |
| `currency` | string or null | Currency code or symbol-derived code |
| `availability` | string or null | Public availability text |
| `colors` | array of strings | Public color names at scrape time |
| `sizes` | array of strings | Public size labels at scrape time |
| `images` | array of strings | Public high-resolution image URLs |
| `category` | string or null | Public category or breadcrumb text |
| `url` | string | Canonical Zara URL of the product page |
| `sourceType` | string | Discovery source such as search term, direct URL, or listing URL |
| `searchTerm` | string or null | Search term used for discovery |
| `scrapedAt` | string | UTC timestamp for validation and save time |

#### Illustrative JSON record

```json
{
  "productId": "43284705",
  "reference": "5644/401/800",
  "name": "LINEN BLEND SHIRT",
  "description": "Relaxed fit shirt made from a linen blend.",
  "price": 49.9,
  "originalPrice": 69.9,
  "currency": "USD",
  "availability": "In stock",
  "colors": ["Ecru", "Navy blue"],
  "sizes": ["S", "M", "L"],
  "images": [
    "https://static.zara.net/photos///2026/V/0/1/p/5644/401/800/2/w/750/5644401800_1_1_1.jpg"
  ],
  "category": "WOMAN / SHIRTS",
  "url": "https://www.zara.com/us/en/linen-blend-shirt-p05644401.html",
  "sourceType": "searchTerm",
  "searchTerm": "linen shirt",
  "scrapedAt": "2026-07-15T17:00:00.000Z"
}
```

### How it works

The Actor uses a browser session with the default Apify Residential proxy configuration. It accepts direct Zara URLs or builds Zara search URLs from search terms and the selected storefront. For search and category pages, it uses listing scrolls to load more product cards up to the configured cap. Each product is validated before it is written to the dataset, and only complete schema-validated product rows are billed as `Product scraped` events. The default dataset contains product rows, while `OUTPUT` and `RUN_SUMMARY` provide terminal outcome and safe diagnostics.

### Pricing

This Actor uses Pay per event pricing. The primary event is `Product scraped`, which is charged only when one complete, schema-validated Zara product record is saved to the default dataset. Apify platform usage also applies, including actor start charges, browser execution, and proxy usage.

For example, twenty-five saved products mean twenty-five product events. The live Pricing tab in Apify Console is the current source of truth for the latest pricing details and platform usage view.

### Use with AI agents (MCP)

This Actor is available as an Apify Actor usable through Apify MCP. The exact Actor identity is `khadinakbar/zara-product-scraper`.

Tool description: use this tool when an agent needs a focused Zara product extraction workflow from product URLs, search terms, or category pages, with validated rows and provenance fields suitable for downstream reasoning.

> Scrape Zara product rows for these search terms and return the validated dataset items with price, availability, images, colors, sizes, canonical URLs, and scrape timestamps. Use the US storefront and cap the run at 8 products.

Output interpretation: the default dataset contains one record per validated product. `sourceType` and `searchTerm` explain how each row was discovered, while `scrapedAt` marks validation time in UTC. Provenance stays inside the product row, so an agent can connect each item back to the search term or direct URL used for discovery.

Scope, pagination, and cost guidance: direct product URLs map to individual product pages, while search and category pages may use multiple listing scrolls. `maxProducts` serves as the global output and billing cap, and `maxScrollsPerListing` controls how many times listing pages load more cards. The default Residential proxy setup matches the browser-driven storefront access pattern.

### Apify API example

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({
  token: process.env.APIFY_TOKEN,
});

const run = await client.actor('khadinakbar/zara-product-scraper').call({
  searchTerms: ['linen shirt'],
  country: 'US',
  maxProducts: 5,
  includeVariants: true,
});

const datasetId = run.defaultDatasetId;
const { items } = await client.dataset(datasetId).listItems({ clean: true });

console.log(items);
```

### Best results and outcome guidance

Use direct product URLs when you already have a known Zara item and want the clearest product-level result. Use `searchTerms` when discovery starts from a concept such as `black dress`, `linen shirt`, or `kids trainers`. Use `startUrls` for product pages, search pages, or category pages that already define the scope. Keep `maxProducts` aligned with the number of product rows you want, and use `includeVariants` when color and size options matter for your downstream workflow.

### Focused standalone workflow

This Actor is designed as a focused standalone workflow.

### Design note

I found that the dataset contract requires `productId`, `name`, `colors`, `sizes`, `images`, `url`, `sourceType`, and `scrapedAt` on every saved record, which makes the output consistently product-shaped and easy to consume.

### FAQ

#### When should I use `startUrls` instead of `searchTerms`?

Use `startUrls` when you already have Zara product, category, or search URLs. Use `searchTerms` when you want the Actor to build Zara search pages from product terms.

#### Which storefront does a search term use?

`searchTerms` uses the storefront selected in `country`. Product and category URLs keep the storefront already present in the URL.

#### How does `maxProducts` affect the run?

`maxProducts` sets the global output and billing cap. The Actor checks it before scheduling products and again before each paid dataset write.

#### What does `includeVariants` add?

It includes publicly visible color and size options from the product page. Availability remains a product-level signal.

#### How should category pages be used?

Category pages work well when a merchandiser wants a bounded slice of a public Zara catalog and expects listing-style discovery with scroll-based loading.

### Responsible use

Use this Actor for public Zara product data in line with applicable terms, law, and organizational policy. The outputs are observations from the public storefront and fit workflows for analysis, catalog monitoring, and agent automation that respect the source site's public access patterns.

# Actor input Schema

## `startUrls` (type: `array`):

Use this when you already have Zara URLs to scrape. Add full Zara product pages, category pages, or search URLs such as https://www.zara.com/us/en/search?searchTerm=linen. Duplicate URLs are ignored and non-Zara domains are rejected. This field is not for product names without a URL.

## `searchTerms` (type: `array`):

Use this when you want the Actor to construct Zara search pages from product terms. Enter words such as linen shirt or black dress; each term is a separate discovery query. The selected country controls constructed search URLs. This field is not needed when startUrls already contains a Zara search page.

## `country` (type: `string`):

Use this when searchTerms should run on a particular Zara storefront. Select US, UK, Spain, Germany, France, Italy, Portugal, Brazil, or Mexico; product and category URLs keep their own storefront. The default US creates https://www.zara.com/us/en/search URLs. This field does not rewrite URLs supplied in startUrls.

## `maxProducts` (type: `integer`):

Use this to set a strict global output and billing cap. The Actor stops scheduling products before the cap and checks it again immediately before every paid dataset write. The default of 25 costs at most $0.10 in product events. This field is not a page-count setting.

## `maxScrollsPerListing` (type: `integer`):

Use this when a Zara search or category page needs more lazy-loaded product cards. Each scroll waits briefly for new cards, then discovery stops at the configured cap. The default of 3 keeps the browser run short and predictable. This field has no effect for direct product URLs.

## `includeVariants` (type: `boolean`):

Use this when you need product color and size options exposed on the product page. Disable it for a smaller record and slightly faster detail extraction. Availability remains a product-level signal and is not inferred from missing size controls. This field does not add cart or checkout automation.

## `proxyConfiguration` (type: `object`):

Use this to override the default Apify Residential proxy configuration when your account has a compatible proxy setup. The default uses the RESIDENTIAL group because Zara applies bot protections. Keep a sticky session per browser request for consistent cookies. This field is not for pasting proxy credentials.

## Actor input object example

```json
{
  "startUrls": [
    "https://www.zara.com/us/en/search?searchTerm=linen"
  ],
  "searchTerms": [
    "linen"
  ],
  "country": "US",
  "maxProducts": 5,
  "maxScrollsPerListing": 1,
  "includeVariants": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Validated product records with price, stock, images, variants, and source URLs.

## `output` (type: `string`):

Compact stable terminal outcome and result count.

## `runSummary` (type: `string`):

Detailed safe diagnostics, cost-cap state, and source counts.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://www.zara.com/us/en/search?searchTerm=linen"
    ],
    "searchTerms": [
        "linen"
    ],
    "country": "US",
    "maxProducts": 5,
    "maxScrollsPerListing": 1,
    "includeVariants": true,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("khadinakbar/zara-product-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["https://www.zara.com/us/en/search?searchTerm=linen"],
    "searchTerms": ["linen"],
    "country": "US",
    "maxProducts": 5,
    "maxScrollsPerListing": 1,
    "includeVariants": True,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("khadinakbar/zara-product-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://www.zara.com/us/en/search?searchTerm=linen"
  ],
  "searchTerms": [
    "linen"
  ],
  "country": "US",
  "maxProducts": 5,
  "maxScrollsPerListing": 1,
  "includeVariants": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call khadinakbar/zara-product-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=khadinakbar/zara-product-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/MakfpK7nvbYrQm7Uc/builds/ncDVxqxiSKZZmpCKQ/openapi.json
