# 🧩 Capterra Scraper (`scraper-engine/capterra-scraper`) Actor

🧩 Capterra Scraper extracts structured software reviews & pricing data from Capterra. 🚀 Great for B2B research, lead gen, competitive analysis & market insights—fast, accurate, and SEO-friendly. 📊✨

- **URL**: https://apify.com/scraper-engine/capterra-scraper.md
- **Developed by:** [Scraper Engine](https://apify.com/scraper-engine) (community)
- **Categories:** Developer tools, Lead generation, Automation
- **Stats:** 1 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.99 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 🧩 Capterra Scraper

Extract clean, structured data from **Capterra** — software products, ratings, pricing, features, deployment & support options, and real client reviews — straight into a tidy dataset you can export to JSON, CSV, or Excel. 🚀

Paste a **product detail** URL or a **category/listing** URL and get back fully normalized records. Reviews can be attached to each product, or pulled on their own in a dedicated reviews-only mode.

***

### ✨ Why Choose This Actor?

- 🧩 **Two URL types in one** — product detail pages *and* category/listing pages.
- 💬 **Rich reviews** — pros/cons, per-criterion ratings, reviewer profile, timestamps and vendor responses.
- 📦 **Bulk friendly** — drop in as many URLs as you like; results are de-duplicated automatically.
- ⚡ **Fast & proxy-free** — fetches over plain HTTP (no browser), works straight out of the box with **no proxy required**.
- 💾 **Live results** — every product/review is saved the moment it's scraped, so you never lose a partial run.

***

### 🎯 Key Features

| Feature | Description |
|--------|-------------|
| 🎯 Product detail | Paste a `https://www.capterra.com/p/<id>/<slug>/` URL to scrape a single product. |
| 📂 Category / listing | Paste a `https://www.capterra.com/<category>-software/` URL to enumerate products across listing pages. |
| 💬 Reviews | Toggle `includeReviews` to attach reviews, or `reviewsOnly` to output reviews alone. |
| 🔢 Max items | Cap the total number of items saved across all URLs. |
| 📄 Pagination | `endPage` (listings) and `endPageForReviews` (reviews) bound how far each goes. |
| 🛡️ Proxy policy | **Optional** — works with no proxy. A direct → datacenter → residential ladder is kept only as a fallback; a custom proxy, if supplied, is tried first. Logged clearly. |

***

### 📥 Input

| Field | Type | Description |
|-------|------|-------------|
| `startUrls` | array | One or more Capterra product/listing URLs. **Required.** |
| `includeReviews` | boolean | Attach reviews to each product. Default `false`. |
| `reviewsOnly` | boolean | Output only review records. Default `false`. |
| `maxItems` | integer | Max items saved across all URLs. Empty = unlimited. |
| `endPage` | integer | Last listing page per URL. Empty = until empty. |
| `endPageForReviews` | integer | Last reviews page per product. Empty = all. |
| `proxyConfiguration` | object | Apify proxy settings powering the fallback ladder. |

#### Example input

```json
{
  "startUrls": [
    "https://www.capterra.com/p/162035/Filestage/",
    "https://www.capterra.com/business-intelligence-software/"
  ],
  "includeReviews": true,
  "reviewsOnly": false,
  "maxItems": 10,
  "endPage": 1,
  "endPageForReviews": 2,
  "proxyConfiguration": { "useApifyProxy": true }
}
```

***

### 📤 Output

Each product record (reviews nested when enabled):

```json
{
  "productId": "162035",
  "name": "Filestage",
  "productUrl": "https://www.capterra.com/p/162035/Filestage/",
  "logoUrl": "https://.../ProductLogo/....png",
  "reviewCount": 102,
  "overallRating": 4.71,
  "easeOfUseRating": 4.7,
  "customerServiceRating": 4.7,
  "valueForMoneyRating": 4.6,
  "longDescription": "Filestage is the online proofing software ...",
  "pricingDetails": {
    "startingPrice": "€199",
    "pricingModel": "Flat Rate",
    "paymentFrequency": "Per Month",
    "hasFreeTrial": true,
    "hasFreeVersion": true
  },
  "features": [{ "title": "Collaboration Tools", "rating": 4.65, "reviewCount": 26 }],
  "training": [{ "value": "5", "label": "Documentation", "selected": true }],
  "support": [{ "value": "1", "label": "Email/Help Desk", "selected": true }],
  "targetCompanySizes": [{ "value": "2", "label": "2-10", "selected": true }],
  "deploymentOptions": [{ "value": "1", "label": "Cloud, SaaS, Web-Based", "selected": true }],
  "category": { "name": "Collaboration Software", "slug": "/collaboration-software" },
  "relatedProducts": [{ "productId": "147657", "name": "monday.com", "overallRating": 4.6 }],
  "reviewUrl": "https://www.capterra.com/p/162035/Filestage/reviews/",
  "reviews": [
    {
      "reviewId": "3080815",
      "title": "The best tool on the market",
      "writtenOn": "September 10, 2021",
      "prosText": "Filestage offers only the features you really need ...",
      "consText": "Due the features are developed for the needs ...",
      "overallRating": "5.0",
      "reviewer": { "fullName": "Erich W.", "jobTitle": "CEO", "companySize": "11-50 employees" }
    }
  ]
}
```

In **reviews-only** mode each row is a single review object.

***

### 🚀 How to Use (Apify Console)

1. Log in at https://console.apify.com → **Actors**.
2. Open the **Capterra Scraper** actor.
3. Paste your Capterra product/listing URLs and configure the toggles (reviews, max items, pagination, proxy).
4. Click **Start**.
5. Watch the logs in real time as products and reviews stream in.
6. Open the **Output** tab — switch between the **🧩 Products** and **⭐ Reviews** table views.
7. Export to JSON / CSV / XLSX.

### 🤖 Use via API

```bash
curl -X POST "https://api.apify.com/v2/acts/<ACTOR_ID>/runs?token=$APIFY_TOKEN" \
     -H "Content-Type: application/json" \
     -d '{"startUrls":["https://www.capterra.com/p/162035/Filestage/"],"includeReviews":true,"maxItems":10}'
```

***

### 💡 Best Use Cases

- 📊 Build software-catalog datasets (pricing, ratings, features) for benchmarking.
- 🗣️ Aggregate client reviews for sentiment and ratings analysis.
- 🧭 Map categories and product coverage for market research.
- 🧪 Enrich product records with per-feature ratings and deployment/support options.

***

### 💵 Pricing

This Actor uses **pay-per-event** billing: you are charged once per item saved to the dataset (`row_result`) — that's one product, or one review in reviews-only mode. You only pay for the data you actually get back.

***

### ❓ FAQ

**Which URLs are supported?** Product detail (`/p/<id>/<slug>/`) and category/listing (`/<category>-software/`) pages. Capterra retired the `/services/*` section, so those URLs are skipped automatically.

**Do I need a proxy?** No. The Actor fetches Capterra over plain HTTP and works directly from Apify's servers with no proxy. The proxy ladder (direct → datacenter → residential) is only a fallback for the rare blocked request; a custom proxy, if supplied, is tried first.

**Will I lose data if the run stops early?** No. Every item is saved live as it's scraped.

***

### 🛟 Support and Feedback

Found a bug or need a tweak? Reach us at **dev.scraperengine@gmail.com** — we'd love to help. 🤝

***

### ⚖️ Legal

Data is collected only from **publicly available** Capterra pages. You are responsible for complying with applicable laws (GDPR, CCPA, etc.) and Capterra's Terms of Service.

# Actor input Schema

## `startUrls` (type: `array`):

📥 Paste one or more Capterra URLs. Supports bulk input — add as many as you like.

✅ Supported URL types:
• 🎯 Product detail — e.g. https://www.capterra.com/p/162035/Filestage/
• 📂 Category / listing — e.g. https://www.capterra.com/business-intelligence-software/

## `includeReviews` (type: `boolean`):

📝 Attach real client reviews to each product record. Increases run time proportional to the number of review pages fetched.

## `reviewsOnly` (type: `boolean`):

📤 Output only review records (one row per review) and nothing else. Great for sentiment / ratings analysis.

## `maxItems` (type: `integer`):

🎚️ How many items to save in total across ALL start URLs — products, or reviews in reviews-only mode. This is the main control for run size: set 100, 200, 500, etc. and the Actor pages through category listings as deep as needed to collect that many. Leave empty for unlimited.

## `endPage` (type: `integer`):

🛑 Optional hard cap on how many pages of each category/listing URL to read (applied per listing URL). Leave empty (recommended) and let 'Max items' decide how deep to page. Only set this if you specifically want to stop a listing after N pages.

## `endPageForReviews` (type: `integer`):

🛑 Last reviews page to read per product. Leave empty to read all review pages.

## `proxyConfiguration` (type: `object`):

🛡️ Optional. This Actor fetches Capterra over plain HTTP and works directly from Apify's servers — no proxy is required. The connection ladder is 🌐 direct → 🛰️ datacenter → 🏠 residential and it sticks to the first rung that works. 🧷 If you paste your own proxy via the "Custom proxies" option, it is tried first. Leave as-is for best speed/cost.

## Actor input object example

```json
{
  "startUrls": [
    "https://www.capterra.com/p/162035/Filestage/",
    "https://www.capterra.com/business-intelligence-software/"
  ],
  "includeReviews": false,
  "reviewsOnly": false,
  "maxItems": 10,
  "endPageForReviews": 2,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://www.capterra.com/p/162035/Filestage/",
        "https://www.capterra.com/business-intelligence-software/"
    ],
    "maxItems": 10,
    "endPageForReviews": 2,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scraper-engine/capterra-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        "https://www.capterra.com/p/162035/Filestage/",
        "https://www.capterra.com/business-intelligence-software/",
    ],
    "maxItems": 10,
    "endPageForReviews": 2,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("scraper-engine/capterra-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://www.capterra.com/p/162035/Filestage/",
    "https://www.capterra.com/business-intelligence-software/"
  ],
  "maxItems": 10,
  "endPageForReviews": 2,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call scraper-engine/capterra-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scraper-engine/capterra-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/aw11rjprk4bbctUhG/builds/7kkw7AeajDdV4UZAo/openapi.json
