# AI Job Search Agent — Open-Web Job Finder (BYO AI Key) (`nomad-agent/web-search-scraper`) Actor

AI agent that scans the open web for job postings beyond any single board. Returns an LLM match score with reasoning, salary, and HTTP-verified live links. $3/1,000 results. Bring your own Anthropic or Mistral key (Mistral path ~$0.01/run).

- **URL**: https://apify.com/nomad-agent/web-search-scraper.md
- **Developed by:** [Nomad.Dev](https://apify.com/nomad-agent) (community)
- **Categories:** Jobs, AI, Agents
- **Stats:** 4 total users, 1 monthly users, 90.2% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.30 / 1,000 job results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## AI Job Search Agent — Web Job Finder (BYOK)

> **Claude / Codex skill to describe and setup this actor: [SKILL.md](https://github.com/Exdenta/OinkAIJobSearch/blob/main/skill/web-search-scraper/SKILL.md)**

An AI agent (BYO API key — Anthropic or Mistral) that hunts the open web for postings matching your query — company career pages, niche boards, anywhere a single job board wouldn't cover.

> **Bring your own key.** Pick a `provider` in the input: `anthropic` (default) uses Claude with its built-in web-search tool; `mistral` uses keenable (no-auth open-web search) for discovery plus a Mistral model for per-page judging/extraction. Leave `provider` unset and just supply whichever key you have — the Actor auto-selects the matching provider. Either way, model usage is **billed separately by the provider you picked**, see [Pricing](#pricing) below. **Without a valid key for the selected provider the run still succeeds** — it produces a single dataset row explaining how to supply a key, instead of failing.

### Two providers, same output

| | `provider: "anthropic"` (default) | `provider: "mistral"` |
|---|---|---|
| Discovery | One Claude agent loop forms its own queries, browses, and judges results (built-in `web_search_20250305` tool) | [keenable](https://keenable.ai) web search (no key needed) — queries built mechanically from your `keywords`/`titleMustMatch` |
| Page judging + extraction | Same agent loop, inline | One Mistral call per candidate page — vetoes index/aggregator/closed/stale pages, extracts the rest |
| Key required | `anthropicApiKey` | `mistralApiKey` |
| Output shape | Identical — same fields, same downstream liveness check either way | |

Pick `mistral` if you'd rather not hold an Anthropic key, or want a cheaper/faster run — testing found it succeeds more reliably per attempt and finishes faster when Claude's agent loop does too, at a small precision cost the extraction prompt corrects for (rejects closed/stale/third-party-repost pages explicitly).

### What this scraper returns

Each result is one flat JSON record per job posting. Before a result is returned, its URL is checked for basic liveness (see [Verification](#verification) below).

| Field | Meaning |
|---|---|
| `id` | Stable identifier derived from the posting URL |
| `source` | Always `"web_search"` for this Actor |
| `title` | Job title as posted |
| `company` | Hiring company / organisation |
| `location` | Location / duty station (may include remote hints) |
| `url` | Direct link to the posting |
| `postedAt` | Posting date where the source provides it (`YYYY-MM-DD`) |
| `snippet` | Short description excerpt |
| `salary` | Salary / compensation text exactly as stated on the posting (e.g. `"$120k–$150k"`, `"Competitive"`), extracted by the discovery/extraction LLM. `null` when the posting states no salary — the agent never guesses a figure |
| `matchScore` | LLM-judged `0–100` rating of how well the posting matches your profile (keywords, titles, locations, remote, seniority, and your free-text `userDescription`). Results are sorted best-match-first. `null` when the model returned no usable score |
| `matchReasoning` | One short sentence from the LLM explaining the `matchScore` |
| `isNew` | `true` only on delta runs (`onlyNewSinceLastRun`), marking a posting not seen in a previous run. Absent on normal runs |
| `verified` | `true` when the liveness check confirmed the URL reachable at run time; `false` when the check couldn't complete (timeout, connection error) but the result was kept (see below) |

### How to scrape AI job search with this Actor

1. Click **Try for free** / **Run** — no login to the target site, no cookies, no proxies to configure.
2. Adjust the input (keyword, filters, `maxItems`) or keep the defaults.
3. Run it and export the dataset as JSON, CSV or Excel, or read it over the [API](https://docs.apify.com/api/v2).

A typical run takes 1–3 minutes (the agent performs up to 10 live web searches). If you lower the run timeout, allow at least 300 seconds so the run isn't killed mid-search.

Run it from your own code:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("nomad-agent/web-search-scraper").call(run_input={"maxItems": 15, "anthropicApiKey": "sk-ant-..."})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["title"], "—", item["company"], item["url"])
```

Or a single HTTP call that runs the Actor and returns items in one response:

```bash
curl -X POST \
  "https://api.apify.com/v2/acts/nomad-agent~web-search-scraper/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"maxItems": 15, "anthropicApiKey": "sk-ant-..."}'
```

### Verification

"Verified" means one thing only: **the posting URL is HTTP-reachable and does not redirect to the site's homepage/root.** After the discovery agent proposes candidate postings, this Actor issues a `HEAD` (falling back to `GET`) request to each URL with a 10-second timeout and modest concurrency. A result is **dropped** only when:

- the final response is `404` or `410` (the posting is gone), or
- the URL redirects to the site's root path (`/` or empty path) — a common pattern for expired listings.

Anything else — a normal `200`, a `403`/`5xx` from a bot-defensive site, a timeout, or a connection error — is **kept**, since those don't reliably distinguish "blocked us" from "actually dead." Results kept because the check couldn't complete (timeout/connection error) carry `verified: false` so you can tell them apart from confirmed-reachable ones. This is a structural HTTP liveness check, not a claim about listing accuracy, freshness, or relevance — that judgment still comes from the discovery agent and any downstream scoring you apply.

Run-level counts (`rawPostings`, `titleExcluded`, `domainExcluded`, `urlsChecked`, `kept`, `dropped`, `deltaSkipped`) are written to the run's key-value store under `VERIFICATION_STATS` and logged at the end of each run.

### Input

| Field | Type | Default | Notes |
|---|---|---|---|
| `provider` | string (select) | `"anthropic"` | `anthropic` (Claude + built-in web search) or `mistral` (keenable web search + Mistral extraction). See [Two providers](#two-providers-same-output) above. |
| `anthropicApiKey` | string | — (required when `provider: "anthropic"`) | Your Anthropic API key (`sk-ant-…`). |
| `model` | string | `"claude-haiku-4-5-20251001"` | Claude model for the discovery agent (`provider: "anthropic"` only). Haiku is fast and inexpensive; switch to Sonnet or Opus for higher-quality extraction on complex pages. |
| `mistralApiKey` | string | — (required when `provider: "mistral"`) | Your Mistral API key. Used only for per-page judging/extraction — the search itself (keenable) needs no key. |
| `mistralModel` | string (select) | `"mistral-medium-latest"` | Mistral model for per-page judging/extraction (`provider: "mistral"` only). Testing found Small matches Medium on quality for this task at ~5x the speed — switch to Small for faster/cheaper runs. |
| `keywords` | array | — | Role or technology keywords the agent should search for (e.g. "frontend", "react", "typescript"). Each entry is one keyword or short phrase. |
| `locations` | array | — | Preferred locations or "remote". Biases results toward these and skips obvious mismatches. |
| `remote` | string (select) | `"any"` | Remote work preference: `any`, `remote-only`, `hybrid`, `on-site`. `provider: "anthropic"` only — the `mistral` path doesn't currently read this field. |
| `seniority` | string (select) | `"any"` | Target seniority level: `any`, `junior`, `mid`, `senior`, `lead`. `provider: "anthropic"` only — the `mistral` path doesn't currently read this field. |
| `titleMustMatch` | array | — | Both providers: preferred/searched title terms. On `mistral`, used to build the keenable search queries when `keywords` is empty. |
| `titleExclude` | array | — | Postings whose title contains any of these terms are dropped — enforced client-side after discovery, both providers. |
| `maxItems` | integer | `15` | Maximum number of job postings to return (1–30). Either provider may return fewer if it can't find enough good candidates, or if some fail the liveness check. |
| `maxAgeHours` | integer | `168` | Preferred maximum age of postings in hours (minimum 24). `provider: "anthropic"` only — the `mistral` path's extraction prompt has its own fixed ~30-day freshness cutoff instead. |
| `userDescription` | string | — | Free-text description of what you are looking for. Primary signal for both discovery and the per-posting `matchScore`/`matchReasoning`. Read by both providers' scoring step. |
| `onlyNewSinceLastRun` | boolean | `false` | Delta / monitoring mode. Only outputs postings not seen in a previous run that also had this flag on; already-seen postings are dropped before push (**not billed**). State is tracked per Actor in a dedicated key-value store, keyed by each posting's `id`. New records carry `isNew: true`. See [Delta mode](#delta-mode--monitoring). |

**Note on `seniority`:** this field used to accept free text and now uses a fixed dropdown (`any` / `junior` / `mid` / `senior` / `lead`). If you have a saved input configuration with an old free-text value that isn't one of these options, it will fail validation on the next run — open the input and re-select from the dropdown.

### Delta mode / monitoring

Turn on `onlyNewSinceLastRun` to only get postings you haven't seen before. Each run records the `id` of every posting it emits (in a dedicated per-Actor key-value store); a later run with the flag on drops any posting whose `id` was already recorded, **before it is pushed or billed**. New postings carry `isNew: true`.

This is the cheapest way to run the Actor on a schedule: pay only for genuinely new matches. State survives across runs and is trimmed to the most recent 50,000 ids. If the state store can't be opened for some reason, the run continues normally without dedup rather than failing.

### Match scoring & ranking

Every posting is scored `0–100` (`matchScore`) by the same LLM that judges it, against your `keywords`, `titleMustMatch`, `locations`, `remote`, `seniority`, and — most importantly — your free-text `userDescription`. A one-sentence `matchReasoning` explains each score. Results in the dataset are sorted **best-match-first**. This is prompt-driven, not a hardcoded rule set, so it weighs role content in context rather than matching exact title tokens.

### Output example

```json
{
  "id": "ws-3f9a1c2b7d",
  "source": "web_search",
  "title": "Computational Linguist",
  "company": "DeepJudge",
  "location": "Zurich / Remote EU",
  "url": "https://deepjudge.ai/careers/computational-linguist",
  "postedAt": "2026-06-25",
  "snippet": "Found on company careers page during agent web search.",
  "salary": "CHF 120,000–150,000 / year",
  "matchScore": 88,
  "matchReasoning": "Strong NLP/linguistics fit, remote-EU friendly, senior level as requested.",
  "verified": true
}
```

### Integrations

Send results straight to Google Sheets, Slack, Make, Zapier or any webhook via [Apify integrations](https://apify.com/integrations) — no code required, or pull the dataset over the [API](https://docs.apify.com/api/v2).

### Pricing

This is a **BYOK** Actor, so a run costs you two things: what Apify charges, and what your own LLM provider charges. Both are listed here — no surprises on your provider bill.

**1. Apify side** — pay per event: **$0.005 per Actor start** and **$0.003 per job returned** (tiered down to $0.0023 on higher Apify plans).
100 jobs ≈ **$0.31**. No subscription, no rental — you pay only for what you fetch.

**2. Your provider key** — billed to you directly, *not* through Apify:

| Provider | What your key pays for | Typical cost per run |
|---|---|---|
| `mistral` | Extraction only. Web search runs through keenable, which needs **no key and costs you nothing**. Your key pays only for the small per-page judging calls. | **~$0.01** |
| `anthropic` | The Claude agent loop **plus** Anthropic's built-in web-search tool, which Anthropic bills at **$10 per 1,000 searches**. This Actor caps the agent at 6 searches per run (~$0.06), plus tokens. | **~$0.06–0.10** |

**Pick `mistral` to keep costs down** — it is roughly 6–10× cheaper per run and needs no paid search tool. Pick `anthropic` when you want Claude's agent loop to form and refine its own queries and are happy to pay for it. See [Anthropic's pricing](https://www.anthropic.com/pricing) and [Mistral's pricing](https://mistral.ai/pricing).

### Use cases

- Long-tail job discovery beyond big boards
- Passive-candidate tooling (find who hires for X)
- Niche-role hunting (rare stacks, rare titles)
- Backfilling gaps in board coverage

### FAQ

**Is it legal to scrape AI job search?**
This Actor reads only publicly available job postings — data any visitor can see without logging in. No personal data behind authentication is touched. Review the target site's terms and your local regulations for your specific use case.

**Do I need an account on the target site?**
No. Postings are discovered via web search and fetched from public pages — no login, cookies or session tokens.

**Does "verified" mean the listing is accurate or still accepting applications?**
No. It means the URL was HTTP-reachable and did not redirect to the site's homepage at check time (see [Verification](#verification)). A posting can still have closed between the check and when you view it.

**How many jobs can I get?**
`maxItems` caps the run between 1 and 30 (default 15). The agent may return fewer than the cap if it can't find enough good candidates or if some fail the liveness check.

**Does it extract salary?**
Yes. The discovery/extraction LLM copies the salary/compensation text exactly as stated on the posting (e.g. `"$120k–$150k / year"`, `"Competitive"`). When a posting states no salary, `salary` is `null` — the agent never guesses or invents a figure.

**Something broken or missing?**
Open an issue on the Actor's **Issues** tab — it is monitored and reliability fixes ship fast.

**Is this Actor useful to you?**
A quick ⭐ review on the Actor's **Reviews** tab helps other search-data users find it — and tells us what to build next.

### Related Actors

- [LinkedIn Jobs Scraper — No Login, No Cookies](https://apify.com/nomad-agent/linkedin-scraper)
- [Research & Academic Jobs Scraper — 10 Sources](https://apify.com/nomad-agent/researcher-bundle)
- [Web Developer Jobs Scraper — 10 Boards in One](https://apify.com/nomad-agent/web-dev-bundle)

***

**From the maker of [Oink](https://github.com/Exdenta/OinkAIJobSearch)** — an open-source, AI-powered job-search bot for Telegram that runs on these Actors. [Try the free bot](https://t.me/job_search_everyday_bot), get a managed instance at [oinkjobsearch.com](https://oinkjobsearch.com), or browse the [full catalog of 50+ Actors](https://apify.com/nomad-agent).

# Actor input Schema

## `provider` (type: `string`):

Which AI provider runs discovery + extraction. <code>anthropic</code> (default) uses a Claude agent with Anthropic's built-in web-search tool. <code>mistral</code> uses keenable (no-auth web search) for discovery plus a Mistral model for per-page judging/extraction -- pick this if you'd rather bring a Mistral key than an Anthropic one. <code>openai</code> uses the same keenable search plus a GPT model for per-page judging/extraction -- pick this if you'd rather bring an OpenAI key.

## `anthropicApiKey` (type: `string`):

Your Anthropic API key (sk-ant-…). Required when provider is <code>anthropic</code>. The actor uses it to run a web-searching AI agent that discovers job postings. Must start with sk-ant- (an OpenAI-style sk-... key will not work here).<br><br><b>Cost note:</b> this is billed to <i>your</i> Anthropic account, on top of what Apify charges. Anthropic bills web search at <b>$10 per 1,000 searches</b> and this actor runs up to 6 searches per run (~$0.06), plus tokens — typically <b>$0.06-0.10 per run</b>. Choose <code>provider=mistral</code> for a much cheaper run: its web search (keenable) is free and only the per-page extraction uses your key.

## `model` (type: `string`):

Claude model to use for the discovery agent (provider=<code>anthropic</code> only). Haiku is fast and inexpensive and is the recommended default; switch to Sonnet 5 or Opus 4.8 for higher-quality extraction on complex pages (both cost more per run against your own key). The 4.5-era models are kept for backwards compatibility with saved tasks.

## `mistralApiKey` (type: `string`):

Your Mistral API key. Required when provider is <code>mistral</code>. Used only for per-page judging/extraction -- the web search itself goes through keenable, which needs no key.

## `mistralModel` (type: `string`):

Mistral model for per-page judging/extraction (provider=<code>mistral</code> only). Medium is the default. Testing found Small matches Medium on quality for this task (extraction is small and well-scoped) at ~5x the speed -- switch to Small if you want faster/cheaper runs.

## `openaiApiKey` (type: `string`):

Your OpenAI API key (sk-...). Required when provider is <code>openai</code>. Used only for per-page judging/extraction -- the web search itself goes through keenable, which needs no key.

## `openaiModel` (type: `string`):

OpenAI model for per-page judging/extraction (provider=<code>openai</code> only). Defaults to <code>gpt-4.1-mini</code> -- cheap, fast and ample for this well-scoped extraction task. Any chat-completions model works, including the gpt-5 family (e.g. <code>gpt-5.4-mini</code>).

## `keywords` (type: `array`):

Role or technology keywords the agent should search for (e.g. <code>frontend</code>, <code>react</code>, <code>typescript</code>). Each entry is one keyword or short phrase.

## `locations` (type: `array`):

Preferred locations or <code>remote</code>. The agent will bias results toward these and skip obvious mismatches.

## `remote` (type: `string`):

Remote work preference communicated to the search agent. Free text — the AI agent interprets it (e.g. any, remote-only, hybrid, on-site).

## `seniority` (type: `string`):

Target seniority level. The agent will prefer postings that match. Free text — the AI agent interprets it (e.g. junior, mid, senior, lead, principal — or any).

## `titleMustMatch` (type: `array`):

The agent will prefer postings whose title contains at least one of these terms.

## `titleExclude` (type: `array`):

The agent will skip postings whose title contains any of these terms.

## `maxItems` (type: `integer`):

Maximum number of job postings to return (1-50; values outside that range are clamped, not rejected). The agent may return fewer if it cannot find enough high-quality matches.

## `maxAgeHours` (type: `integer`):

Preferred maximum age of postings in hours (default 168 = one week). The agent will prefer fresh postings but may include older ones when fresh results are scarce.

## `userDescription` (type: `string`):

Free-text description of what you are looking for. The agent uses this as primary signal alongside the structured filters above — and it is also the main signal the per-posting <code>matchScore</code>/<code>matchReasoning</code> is judged against.

## `onlyNewSinceLastRun` (type: `boolean`):

Delta / monitoring mode: only output postings not already seen in a previous run that also had this flag on. Already-seen postings are dropped before push (and <b>not billed</b>), so this is the cheapest way to run this Actor on a schedule and pay only for genuinely new postings. State is tracked per Actor in a dedicated key-value store, keyed by each posting's <code>id</code> (a SHA1 hash of its URL). New records carry <code>isNew: true</code>.

## Actor input object example

```json
{
  "provider": "anthropic",
  "model": "claude-haiku-4-5-20251001",
  "mistralModel": "mistral-medium-latest",
  "openaiModel": "gpt-4.1-mini",
  "keywords": [
    "frontend",
    "react",
    "typescript"
  ],
  "locations": [
    "remote",
    "Europe",
    "Spain"
  ],
  "remote": "any",
  "seniority": "any",
  "titleMustMatch": [
    "frontend",
    "react",
    "typescript"
  ],
  "titleExclude": [
    "intern",
    "manager"
  ],
  "maxItems": 15,
  "maxAgeHours": 168,
  "userDescription": "Looking for a senior React engineer role at a product company, ideally in fintech or developer tooling, fully remote within Europe.",
  "onlyNewSinceLastRun": false
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("nomad-agent/web-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("nomad-agent/web-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call nomad-agent/web-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=nomad-agent/web-search-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Jima07w7QhOtcSmRW/builds/pjX0Jkb0RTU6c0mfR/openapi.json
