# Cars.com Vehicles Scraper (`fmchisti/cars-com-scraper`) Actor

Scrape Cars.com shopping results and vehicle detail pages from start URLs.

- **URL**: https://apify.com/fmchisti/cars-com-scraper.md
- **Developed by:** [Fahim Mahmud Chisti](https://apify.com/fmchisti) (community)
- **Categories:** Automation, Integrations, Other
- **Stats:** 3 total users, 3 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

### What does Cars.com Vehicles Scraper do?

**Cars.com Vehicles Scraper** extracts used and new vehicle listings from [Cars.com](https://www.cars.com) shopping search and detail pages. Paste a filtered search URL (zip, radius, mileage, seller type, sort) or a direct vehicle listing URL, and the Actor returns structured fields that match the other vehicle scrapers in this monorepo—title, price, mileage, VIN, seller, location, images, and more.

Run it on the [Apify](https://apify.com) platform for scheduling, API access, proxy rotation, monitoring, and dataset exports (JSON, CSV, Excel, HTML).

### Why use Cars.com Vehicles Scraper?

- Track private-seller and dealer inventory with your Cars.com filters preserved in the start URL
- Monitor listed dates and prices for market research
- Feed CRM tools, marketplaces, or valuation pipelines with consistent vehicle fields
- Paginate search results and optionally open each listing for full specs

### How to use Cars.com Vehicles Scraper

1. Open the Actor in Apify Console and go to the **Input** tab.
2. Paste one or more Cars.com **Start URLs** (shopping search or vehicle detail).
3. Set **Max items** and **Max pages per search** (defaults are low for quick QA runs).
4. Keep **Scrape item details** enabled when you need VIN, engine, colors, and description.
5. Keep US residential proxies enabled if datacenter IPs are blocked.
6. Click **Start** and download results from the **Output** / dataset tab.

### Input

| Field                     | Description                                                                                                 |
| ------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `startUrls`               | Required. Cars.com shopping search or vehicle detail URLs. Filters in the URL are kept.                     |
| `maxItems`                | Cap on saved listings (`0` = unlimited). Default `5`.                                                       |
| `maxPagesPerSearch`       | Pages to crawl per search URL. Default `1`.                                                                 |
| `maxListingAgeDays`       | Optional age filter using the listing listed date.                                                          |
| `scrapeItemDetails`       | Open each vehicle page for seller notes, engine, transmission, interior color, more images. Default `true`. |
| `duplicateCheck`          | Skip detail scraping for listing URLs already known (default `false`).                                      |
| `duplicateCheckApiUrl`    | Optional POST exists API; leave empty for Apify storage only.                                               |
| `duplicateCheckStoreName` | Named KV store (default `vehicle-listing-urls`).                                                            |
| `proxyConfiguration`      | Prefer Apify residential US proxies.                                                                        |

Example input:

```json
{
    "startUrls": [
        {
            "url": "https://www.cars.com/shopping/results/?mileage_max=70000&stock_type=used&seller_type%5B%5D=private_seller&zip=19104&maximum_distance=9999&sort=listed_at_desc"
        }
    ],
    "maxItems": 5,
    "maxPagesPerSearch": 1,
    "scrapeItemDetails": true
}
```

### Skip existing listings (duplicate check)

Cars.com detail scraping opens each vehicle in a browser with residential proxies. On scheduled re-runs, most results are often listings you already stored. Duplicate check skips those URLs so you do not pay again for detail pages you already have.

Enable it with `duplicateCheck: true` (default `false`).

If N consecutive listing URLs in a search are known duplicates (N = `duplicateCheckLeadingStop`, default **20**), that search stops (no further pages) to avoid paying for known inventory.

**Modes**

- When **`duplicateCheckApiUrl`** is set, the Actor POSTs batches of up to 500 listing URLs to your endpoint, then skips URLs returned in `existing`. It also reads and updates the named Apify Key-Value store. A URL is skipped if **either** your API or the store marks it as known.
- When **`duplicateCheckApiUrl`** is empty, the Actor uses the **Apify Key-Value store only** (no external API).

**API contract**

```json
// Request
{ "listingUrls": ["https://example.com/listing/1", "https://example.com/listing/2"] }

// Response
{
  "existing": ["https://example.com/listing/1"],
  "missing": ["https://example.com/listing/2"]
}
```

`listingUrls` may also be a single string. URLs are normalized (query string and trailing slash ignored). No auth header is required for public endpoints. If the API call fails or exceeds a **20s timeout**, the Actor fail-opens and scrapes the batch.

After a listing is saved successfully, its URL is written to the named store under the `KNOWN_LISTING_URLS` record so future runs skip it even without an API. Cached URLs expire after **30 days**, and the store is capped at **50,000** entries (oldest first). Legacy boolean `true` entries are migrated to timestamps on load.

Example:

```json
{
    "startUrls": [{ "url": "https://www.cars.com/shopping/results/?zip=19104&sort=listed_at_desc" }],
    "maxItems": 100,
    "duplicateCheck": true,
    "duplicateCheckApiUrl": "https://your-api.example.com/listings/exists",
    "duplicateCheckStoreName": "vehicle-listing-urls"
}
```

### Output

Each dataset item uses the shared vehicle listing shape. You can download the dataset as JSON, HTML, CSV, or Excel.

```json
{
    "itemId": "example-listing-id",
    "title": "2018 Honda Civic EX",
    "listedAt": "2026-07-20T00:00:00.000Z",
    "price": "18500",
    "currency": "USD",
    "location": "Philadelphia, PA",
    "seller": "Private Seller",
    "year": "2018",
    "make": "Honda",
    "model": "Civic",
    "mileage": "42000",
    "transmission": "Automatic",
    "vin": null,
    "url": "https://www.cars.com/vehicledetail/example-listing-id/",
    "imageUrl": "https://platform.cstatic-images.com/example.jpg",
    "images": ["https://platform.cstatic-images.com/example.jpg"],
    "page": 1,
    "scrapedAt": "2026-07-20T17:00:00.000Z"
}
```

When the Actor is started from an Apify Task, each dataset item also includes `taskId` and `taskName` so you can track which Task produced it. Direct Actor runs set both to `null`.

With `scrapeItemDetails: true`, expect additional fields such as `vin`, `engine`, `bodyType`, `driveType`, exterior/interior colors, `description`, and `specifics`.

### Data fields

| Field                        | Notes                                                     |
| ---------------------------- | --------------------------------------------------------- |
| `itemId`                     | Cars.com listing identifier from the detail URL           |
| `title`                      | Listing title                                             |
| `listedAt`                   | From Cars.com SSR fingerprint / days-on-market when shown |
| `price` / `currency`         | Asking price in USD when present                          |
| `mileage`                    | Normalized odometer when available                        |
| `year` / `make` / `model`    | Parsed from title and detail specs                        |
| `seller` / `forSaleBy`       | e.g. Private Seller or dealer name                        |
| `vin`                        | From detail specs when detail scraping is on              |
| `images` / `imageUrl`        | Cars.com CDN vehicle images                               |
| `url` / `sourceUrl` / `page` | Provenance and pagination                                 |

### Tips

- Build the search on Cars.com first (filters, zip, sort), then copy the URL into `startUrls`.
- Keep `maxItems` and `maxPagesPerSearch` small while testing; raise them for production runs.
- Keep `scrapeItemDetails` enabled for engine, transmission, interior color, seller notes, and gallery images—each listing costs an extra request. Search cards already include VIN and core specs.
- Enable `duplicateCheck` on scheduled re-runs to skip listings you already stored.
- Set `scrapeItemDetails: false` when search-card fields are enough.
- Prefer skipping known listings over raising memory.
- Keep browser concurrency low.
- If results are empty, check the run key-value store for `DEBUG_PAGE_*` HTML snapshots.

### FAQ and support

This Actor is intended for legitimate research and inventory monitoring. Respect Cars.com terms of service and applicable laws. For bugs or feature requests, use the Actor **Issues** tab on Apify, or ask for a custom solution if you need marketplace-specific enrichments.

# Actor input Schema

## `startUrls` (type: `array`):

Paste the exact Cars.com shopping or vehicle detail URL from your browser. Filters (mileage, seller type, zip, sort, page) are fetched as-is. When provided, only these URLs are crawled.

## `maxItems` (type: `integer`):

Maximum vehicle listings to save across all start URLs. Set to 0 for no item limit.

## `maxPagesPerSearch` (type: `integer`):

Maximum Cars.com result pages to process for each search start URL.

## `maxListingAgeDays` (type: `integer`):

Optional. Save only listings with a known listed date within this many days. Listings without a date are excluded.

## `scrapeItemDetails` (type: `boolean`):

Open each vehicle detail page for seller notes, engine, transmission, interior color, and more images. Recommended for full field coverage.

## `duplicateCheck` (type: `boolean`):

When enabled, skip detail scraping for listing URLs already known from your duplicate-check API and/or a named Apify Key-Value store. Saves proxy and compute cost on re-runs.

## `duplicateCheckApiUrl` (type: `string`):

Optional. POST endpoint that accepts { "listingUrls": string|string\[] } and returns { "existing": string\[], "missing": string\[] }. Leave empty to use Apify Key-Value store only.

## `duplicateCheckStoreName` (type: `string`):

Named Apify Key-Value store that remembers listing URLs across runs. Used whenever Skip existing listings is enabled.

## `duplicateCheckLeadingStop` (type: `integer`):

When Skip existing listings is enabled, stop that search after this many consecutive listing URLs are known duplicates (in a row). Default 20. Lower to stop sooner; raise to keep scanning longer.

## `proxyConfiguration` (type: `object`):

Proxy settings. Residential US proxies are recommended when Cars.com blocks datacenter traffic.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.cars.com/shopping/results/?mileage_max=70000&stock_type=used&seller_type[]=private_seller&zip=19104&maximum_distance=9999&sort=listed_at_desc"
    }
  ],
  "maxItems": 5,
  "maxPagesPerSearch": 1,
  "scrapeItemDetails": true,
  "duplicateCheck": false,
  "duplicateCheckApiUrl": "https://ccscraperapi.up.railway.app/api/listings/exists",
  "duplicateCheckStoreName": "vehicle-listing-urls",
  "duplicateCheckLeadingStop": 20,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `scrapeState` (type: `string`):

No description

## `debugItems` (type: `string`):

No description

## `debugPages` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.cars.com/shopping/results/?mileage_max=70000&stock_type=used&seller_type[]=private_seller&zip=19104&maximum_distance=9999&sort=listed_at_desc"
        }
    ],
    "maxItems": 5,
    "maxPagesPerSearch": 1,
    "scrapeItemDetails": true,
    "duplicateCheckApiUrl": "https://ccscraperapi.up.railway.app/api/listings/exists",
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "US"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("fmchisti/cars-com-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://www.cars.com/shopping/results/?mileage_max=70000&stock_type=used&seller_type[]=private_seller&zip=19104&maximum_distance=9999&sort=listed_at_desc" }],
    "maxItems": 5,
    "maxPagesPerSearch": 1,
    "scrapeItemDetails": True,
    "duplicateCheckApiUrl": "https://ccscraperapi.up.railway.app/api/listings/exists",
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "US",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("fmchisti/cars-com-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.cars.com/shopping/results/?mileage_max=70000&stock_type=used&seller_type[]=private_seller&zip=19104&maximum_distance=9999&sort=listed_at_desc"
    }
  ],
  "maxItems": 5,
  "maxPagesPerSearch": 1,
  "scrapeItemDetails": true,
  "duplicateCheckApiUrl": "https://ccscraperapi.up.railway.app/api/listings/exists",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}' |
apify call fmchisti/cars-com-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=fmchisti/cars-com-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/SQj3Yt9XdHeBLGtef/builds/PFLLHFgf5TClGsIRN/openapi.json
