# Website Change Monitor (`foo121/website-change-monitor`) Actor

Monitor any list of URLs for true content changes — CSS-selector targeting, keyword watches, noise-filtered diffing and webhook-on-change. Built for scheduled runs. Pay per result.

- **URL**: https://apify.com/foo121/website-change-monitor.md
- **Developed by:** [ziv shay](https://apify.com/foo121) (community)
- **Categories:** Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$5.00 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Website Change Monitor

**Monitor any list of URLs for *true* content changes — and fire a webhook only when something real changes.** CSS-selector targeting, keyword watches, noise-filtered diffing, and clean machine-consumable output. Built for scheduled runs. Pay per result.

Most "website change detection" actors diff raw HTML, so they fire on every page load — cache busters, session tokens, ad slots and timestamps all look like "changes." This actor normalizes the page first (strips scripts/styles/comments, volatile attributes, hashes, session ids, cache-buster query params, whitespace reflow) and compares against the **last scheduled run**, so you get signal, not noise.

### What it does

- Fetches each URL, extracts the watched region (whole page or a **CSS selector**).
- Normalizes content and compares it to the previous run's snapshot (stored in a persistent named Key-Value store — it survives between scheduled runs).
- Emits a structured **change-event** per URL that actually changed: added/removed lines, counts, a unified-diff preview, a 0–1 change ratio, and keyword appeared/disappeared signals.
- A `minChangeRatio` threshold suppresses trivial edits.

### Why it's different

- **True-change detection, not raw-HTML diffing.** Volatile noise is stripped before comparison, so you don't get paged on a cache-buster.
- **State persists across scheduled runs.** It compares against your *last* run, not within one run — the only way change detection actually works on a schedule.
- **Clean output.** `addedLines` / `removedLines` / `changeRatio` / `keywordsAppeared` are webhook-ready — your consumer acts on the fields, no HTML parsing.

### Input

| Field | Type | Description |
|---|---|---|
| `startUrls` | array | URLs to watch. Strings, or objects with per-URL `url` + `selector` + `keywords`. |
| `selector` | string | Global CSS selector applied to URLs that don't set their own. Blank = whole `<body>`. |
| `keywords` | array | Words to watch; run reports present + appeared/disappeared (e.g. `In stock`, `Sold out`). |
| `mode` | string | `text` (default, ignores markup churn) or `html`. |
| `minChangeRatio` | number | 0–1. Only flag a change when ≥ this fraction of lines changed. |
| `emitUnchanged` | boolean | Also output unchanged URLs (default off → webhook fires only on real changes). |
| `storeArchive` | boolean | Store previous content for line-level diffs (default on). |
| `proxyConfiguration` | object | Apify proxy (recommended). |

#### Example input

```json
{
  "startUrls": [
    "https://news.ycombinator.com",
    { "url": "https://shop.example.com/item", "selector": ".price", "keywords": ["In stock", "Sold out"] }
  ],
  "minChangeRatio": 0.02
}
```

### Output

One change-event per changed/first-seen/errored URL:

```json
{
  "url": "https://shop.example.com/item",
  "changed": true,
  "changeType": "changed",
  "fingerprint": "a1b2c3d4e5f60718",
  "previousFingerprint": "9988776655443322",
  "addedLines": ["$24.99", "In stock"],
  "removedLines": ["$29.99", "Sold out"],
  "addedCount": 2,
  "removedCount": 2,
  "changeRatio": 0.5,
  "keywordsAppeared": ["In stock"],
  "keywordsDisappeared": ["Sold out"],
  "preview": "- $29.99\n- Sold out\n+ $24.99\n+ In stock",
  "checkedAt": "2026-06-21T13:40:00.000Z",
  "previousCheckedAt": "2026-06-21T12:40:00.000Z"
}
```

`changeType` is one of: `first_seen`, `unchanged`, `below_threshold`, `changed`, `error`.

### Schedule + webhook

1. **Schedule** the actor (e.g. hourly/daily) — the named state store remembers the last snapshot between runs.
2. Add an Apify **webhook** on *run succeeded* and, in your endpoint, act on rows where `changed: true`. Or use the dataset's `changed`/`changeType` fields to filter.

### Pricing

Pay per result — you're charged per change-event row returned. Default output mode only returns rows that changed (or were first-seen / errored), so steady-state monitoring of stable pages is cheap.

***

*Built by MoneyMachine. Public data only; respect each site's terms and robots.*

# Actor input Schema

## `startUrls` (type: `array`):

List of pages to watch for changes. Paste plain URLs, or objects to set a per-URL CSS `selector` and `keywords`, e.g. { "url": "https://shop.com/item", "selector": ".price", "keywords": \["In stock", "Sold out"] }. Each URL produces at most one change-event row per run.

## `selector` (type: `string`):

Optional CSS selector applied to every URL that doesn't set its own. Narrows watching to one region (e.g. ".price", "main", "#content") so changes elsewhere on the page are ignored. Leave blank to watch the whole <body>.

## `keywords` (type: `array`):

Optional keywords. Each run reports which are present, plus which appeared/disappeared since last run (e.g. watch "In stock" / "Sold out" / "Apply now").

## `mode` (type: `string`):

Compare visible text (recommended — ignores markup churn) or raw HTML of the selected region.

## `minChangeRatio` (type: `number`):

Suppress trivial edits: a change is only flagged when the fraction of changed lines is at least this value. 0 = report any change. 0.05 = ignore <5% edits.

## `emitUnchanged` (type: `boolean`):

If on, every checked URL produces a row (changed or not). Default only outputs rows that changed, were first-seen, or errored — so a webhook fires only on real changes.

## `storeArchive` (type: `boolean`):

Store the previous normalized content so the next run can produce added/removed line diffs. Turn off to compare by fingerprint only (lighter, no diff preview).

## `requestTimeoutSecs` (type: `integer`):

Abort a single fetch after this many seconds.

## `maxConcurrency` (type: `integer`):

How many URLs to check in parallel.

## `proxyConfiguration` (type: `object`):

Proxy for outbound requests. Apify proxy recommended to avoid rate limits across scheduled runs.

## Actor input object example

```json
{
  "startUrls": [
    "https://news.ycombinator.com",
    {
      "url": "https://www.iana.org/domains/reserved",
      "selector": "#main_right"
    }
  ],
  "selector": "",
  "keywords": [],
  "mode": "text",
  "minChangeRatio": 0,
  "emitUnchanged": false,
  "storeArchive": true,
  "requestTimeoutSecs": 30,
  "maxConcurrency": 10,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://news.ycombinator.com",
        {
            "url": "https://www.iana.org/domains/reserved",
            "selector": "#main_right"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("foo121/website-change-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [
        "https://news.ycombinator.com",
        {
            "url": "https://www.iana.org/domains/reserved",
            "selector": "#main_right",
        },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("foo121/website-change-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://news.ycombinator.com",
    {
      "url": "https://www.iana.org/domains/reserved",
      "selector": "#main_right"
    }
  ]
}' |
apify call foo121/website-change-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=foo121/website-change-monitor",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/BS18NM84hdxi9JItd/builds/89DLO7UlK1fWOxGMc/openapi.json
