# InfoJobs Scraper — Ofertas de Trabajo y Empleo España (`nomad-agent/infojobs-scraper`) Actor

Scrape ofertas de trabajo y empleo from InfoJobs.net, Spain's largest job board. Keyword, province, contract, jornada and teletrabajo filters; AI skill tags, structured salary (min/max), postedAt and apply URL. Delta mode for alert bots. For recruiters, job boards and Spanish market analysis.

- **URL**: https://apify.com/nomad-agent/infojobs-scraper.md
- **Developed by:** [Nomad.Dev](https://apify.com/nomad-agent) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 3 total users, 3 monthly users, 97.8% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 job results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## InfoJobs Scraper — Ofertas de Trabajo y Empleo España

> **Claude / Codex skill to describe and setup this actor: [SKILL.md](https://github.com/Exdenta/OinkAIJobSearch/blob/main/skill/infojobs-scraper/SKILL.md)**

Scrape current ofertas de trabajo y empleo from InfoJobs.net, the biggest job board in Spain, with keyword and province filtering — every posting enriched with AI-extracted skill tags and structured salary data.

### What InfoJobs data does this scraper extract?

Each result is one flat JSON record per job posting:

| Field | Meaning |
|---|---|
| `title` | Job title as posted |
| `company` | Hiring company / organisation |
| `location` | Location / duty station (may include remote hints) |
| `teleworking` | Work mode shown on the card — `Presencial`, `Híbrido` or `Teletrabajo` (`null` when the card omits it) |
| `contractType` | Contract-type label (e.g. `Contrato indefinido`, `Contrato de duración determinada`) |
| `workday` | Workday / jornada label (e.g. `Jornada completa`, `Jornada parcial`) |
| `url` | Direct link to the posting |
| `postedAt` | Posting date/time, computed server-side as ISO-8601 (`YYYY-MM-DD`, or a full timestamp when the source text is minute/hour-precision, e.g. `"Hace 14m"`) from InfoJobs' relative Spanish text. `null` when that text doesn't match a known pattern. |
| `postedAtText` | The original relative Spanish text as shown by the source (e.g. `"Hace 2h"`, `"Hace 6d"`), kept **verbatim** for transparency (never rewritten, never dropped) |
| `salary` | Salary text where the source provides it |
| `salaryMin` / `salaryMax` | Numeric pay bounds parsed from the salary text (equal for a single figure; `null` when no salary is published) |
| `salaryCurrency` | Currency code for the salary bounds (defaults to `EUR`) |
| `salaryPeriod` | Pay period — `year`, `month`, `week`, `day` or `hour` |
| `skills` | AI-extracted skill tags (max 8) — concrete skills, technologies and tools from the title/description, canonicalised to English industry terms (e.g. `Python`, `SAP`, `Customer service`). Included in the result price; `null` if disabled via `extractSkills: false` or unavailable |
| `snippet` | Short description excerpt |
| `isNew` | `true` for postings new to delta mode (`onlyNewSinceLastRun`); absent on ordinary runs |
| `id` | Stable source-side identifier |
| `source` | Always `"infojobs"` — handy when merging datasets from several job scrapers |

**Transparent posting dates:** InfoJobs shows only relative Spanish text (`Hace 2h`, `Hace 6d`). We keep that text **verbatim** in `postedAtText` and expose the computed ISO date in `postedAt` — unparseable dates are never silently dropped, so you can always audit how each date was derived.

### How to scrape InfoJobs with this Actor

1. Click **Try for free** / **Run** — no login to the target site, no cookies, no proxies to configure.
2. Adjust the input (keyword, filters, `maxItems`) or keep the defaults.
3. Run it and export the dataset as JSON, CSV or Excel, or read it over the [API](https://docs.apify.com/api/v2).

Run it from your own code:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("nomad-agent/infojobs-scraper").call(run_input={"maxItems": 50})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["title"], "—", item["company"], item["url"])
```

Or a single HTTP call that runs the Actor and returns items in one response:

```bash
curl -X POST \
  "https://api.apify.com/v2/acts/nomad-agent~infojobs-scraper/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"maxItems": 50}'
```

### Input

| Field | Type | Default | Notes |
|---|---|---|---|
| `keyword` | string | `""` | Free-text search query (e.g. "data engineer", "marketing"). Leave empty to return the latest postings across all categories. |
| `province` | string (select) | `""` | Restrict results to a single Spanish province (e.g. Madrid, Barcelona). Leave as "All provinces" to search nationwide. |
| `teleworking` | string | `""` | Keep only `Presencial`, `Híbrido` or `Teletrabajo` listings (accent/case-insensitive match on the card's work-mode label). |
| `contractType` | string | `""` | Keep only listings matching a contract-type label (Indefinido, De duración determinada, Fijo discontinuo, …). |
| `workday` | string | `""` | Keep only listings matching a jornada label (Completa, Parcial, Intensiva, Indiferente). |
| `postedSince` | integer | `0` | Only return listings posted within this many days. See "Freshness filtering" below. Set 0 to disable. |
| `titleExclude` | array of strings | `[]` | Skip listings whose title contains any of these case-insensitive terms. |
| `companyExclude` | array of strings | `[]` | Skip listings whose company name contains any of these case-insensitive terms. |
| `extractSkills` | boolean | `true` | Add an AI-extracted `skills` tag array to each posting (included in the result price). Set `false` to skip. |
| `onlyNewSinceLastRun` | boolean | `false` | Delta / monitoring mode: return only postings not seen in a previous run and stamp each with `isNew: true`. See "Delta mode" below. |
| `maxItems` | integer | `15` | Maximum number of postings to return. InfoJobs returns roughly 10–15 cards per page; set 0 for no additional cap. |
| `timeoutSecs` | integer | `25` | How long to wait for InfoJobs to respond before giving up, in seconds. *(Advanced)* |
| `cacheTtlSeconds` | integer | `1800` | Reuses the last fetch for this many seconds so rapid re-runs don't hit InfoJobs again. Set 0 to always fetch live. *(Advanced)* |

#### Freshness filtering (`postedSince`)

InfoJobs shows posting dates as relative Spanish text on the search page, not an ISO date — typically a compact badge like `"Hace 14m"` (minutes), `"Hace 2h"` (hours) or `"Hace 6d"` (days), occasionally a spelled-out form like `"Hace 2 semanas"` or `"Hace 3 meses"`. This text is parsed into an approximate age in days both to evaluate `postedSince` and to compute the ISO `postedAt` field in the output; the original, unmodified text is always kept in `postedAtText`. Listings whose date text doesn't match a known pattern are always kept by the `postedSince` filter, since their age can't be judged — `postedAt` is `null` for those.

#### Delta mode (`onlyNewSinceLastRun`)

For recurring runs (e.g. a daily job-alert bot), set `onlyNewSinceLastRun: true`. The Actor remembers every offer ID it has already returned in a named key-value store and, on the next run, emits **only postings it hasn't seen before**, each stamped with `isNew: true`. This gives you a clean incremental feed with no client-side de-dup. Keep the rest of your input stable across scheduled runs so the seen-set stays comparable; the store holds up to 50,000 IDs (oldest evicted first).

### Output example

```json
{
  "id": "9982314",
  "title": "Desarrollador/a Full Stack",
  "company": "Indra",
  "location": "Madrid, Spain, Híbrido",
  "teleworking": "Híbrido",
  "contractType": "Contrato indefinido",
  "workday": "Jornada completa",
  "url": "https://www.infojobs.net/madrid/desarrollador-full-stack/of-i9982314",
  "postedAt": "2026-07-04T10:00:00+00:00",
  "postedAtText": "Hace 2h",
  "salary": "35.000€ - 45.000€ Bruto/año",
  "salaryMin": 35000,
  "salaryMax": 45000,
  "salaryCurrency": "EUR",
  "salaryPeriod": "year",
  "skills": ["JavaScript", "React", "Node.js", "SQL"],
  "snippet": "Buscamos desarrollador/a full stack...",
  "source": "infojobs"
}
```

### Integrations

Export results as JSON, CSV or Excel; connect via Make, Zapier or n8n; call directly with `run-sync-get-dataset-items`; or plug into AI agents through the Apify MCP server.

### Pricing

Pay per event: **$0.005 per Actor start** and **$0.0025 per job returned** (less on higher Apify plans).
100 jobs ≈ $0.26. No subscription, no rental — you pay only for what you fetch.

### Use cases

- Spanish job boards and alert bots (empleo en España)
- Recruiting agencies sourcing Spain-based talent
- Salary benchmarking for the Spanish market
- Regional labour-market dashboards

### FAQ

**Is it legal to scrape InfoJobs?**
This Actor reads only publicly available job postings — data any visitor can see without logging in. No personal data behind authentication is touched. Review the target site's terms and your local regulations for your specific use case.

**Do I need an account on the target site?**
No. Postings are fetched from public pages/APIs — no login, cookies or session tokens.

**How fresh is the data?**
Every run fetches live listings. Results are cached for `cacheTtlSeconds` (default 30 min, set 0 to always hit the source live).

**How many jobs can I get?**
`maxItems` caps the run (set 0 where supported for no cap). Most sources paginate from newest to oldest.

**Something broken or missing?**
Open an issue on the Actor's **Issues** tab — it is monitored and reliability fixes ship fast.

**Is this Actor useful to you?**
A short review on the Actor's **Reviews** tab helps other users find it — it takes a minute and is genuinely appreciated.

### Related Actors

- [Europe Jobs Scraper — 14 Sources in One](https://apify.com/nomad-agent/europe-jobs-bundle)
- [Web Developer Jobs Scraper — 10 Boards in One](https://apify.com/nomad-agent/web-dev-bundle)
- [Tecnoempleo Scraper — Spain IT & Tech Jobs](https://apify.com/nomad-agent/tecnoempleo-scraper)
- [EURES Job Scraper — EU Job Mobility Portal](https://apify.com/nomad-agent/eures-scraper)

***

**From the maker of [Oink](https://github.com/Exdenta/OinkAIJobSearch)** — an open-source, AI-powered job-search bot for Telegram that runs on these Actors. [Try the free bot](https://t.me/job_search_everyday_bot), get a managed instance at [oinkjobsearch.com](https://oinkjobsearch.com), or browse the [full catalog of 50+ Actors](https://apify.com/nomad-agent).

# Actor input Schema

## `keyword` (type: `string`):

Free-text search query (e.g. <code>data engineer</code>, <code>marketing</code>). Leave empty to return the latest postings across all categories.

## `province` (type: `string`):

Restrict results to a single Spanish province. Leave as <code>All provinces</code> to search nationwide.

## `teleworking` (type: `string`):

Keep only listings whose work-mode label matches (accent- and case-insensitive substring). Values shown on InfoJobs cards: <code>Presencial</code> (on-site), <code>Híbrido</code> (hybrid), <code>Teletrabajo</code> (remote). Leave empty for all.

## `contractType` (type: `string`):

Keep only listings whose contract-type label matches (accent- and case-insensitive substring of the label shown on the card). Common values: <code>Indefinido</code>, <code>De duración determinada</code>, <code>Fijo discontinuo</code>, <code>Formativo</code>, <code>De relevo</code>, <code>Autónomo</code>, <code>A tiempo parcial</code>, <code>Otros contratos</code>. Leave empty for all.

## `workday` (type: `string`):

Keep only listings whose workday label matches (accent- and case-insensitive substring of the label shown on the card). Common values: <code>Completa</code>, <code>Parcial</code>, <code>Intensiva</code>, <code>Indiferente</code>. Leave empty for all.

## `postedSince` (type: `integer`):

Only return listings posted within this many days. InfoJobs shows posting dates as relative Spanish text (e.g. <code>Hace 14m</code>, <code>Hace 2h</code>, <code>Hace 6d</code>, or spelled out as <code>Hace 2 semanas</code>), which is parsed into an approximate age for this filter. The original text is kept as-is in the `postedAtText` output field; `postedAt` holds the computed ISO date/timestamp. Listings whose date text doesn't match a known pattern are always kept (can't be judged, so they're not dropped). Set 0 (default) to disable and return all fetched listings regardless of age.

## `titleExclude` (type: `array`):

Skip listings whose title contains any of these case-insensitive terms.

## `companyExclude` (type: `array`):

Skip listings whose company name contains any of these case-insensitive terms.

## `extractSkills` (type: `boolean`):

Adds a <code>skills</code> array to each posting — concrete skills, technologies and tools extracted from the title and description excerpt by an LLM and canonicalised to English industry terms (e.g. <code>Python</code>, <code>SAP</code>, <code>Forklift operation</code>). Included in the result price. Fail-open: if extraction is unavailable, <code>skills</code> is null and the run still succeeds.

## `onlyNewSinceLastRun` (type: `boolean`):

Monitoring / delta mode for recurring alert-bot runs: return only postings not seen in a previous run of this Actor, and stamp each with <code>isNew: true</code>. Already-seen offer IDs are remembered in a named key-value store between runs (oldest evicted past 50,000). Keep the same input across scheduled runs so the seen-set stays comparable. Default off.

## `maxItems` (type: `integer`):

Maximum number of postings to return. InfoJobs returns roughly 10–15 cards per page; set 0 for no additional cap.

## `timeoutSecs` (type: `integer`):

How long to wait for InfoJobs to respond before giving up, in seconds.

## `cacheTtlSeconds` (type: `integer`):

Reuses the last fetch for this many seconds so rapid re-runs don't hit InfoJobs again. Set 0 to always fetch live.

## Actor input object example

```json
{
  "keyword": "software engineer",
  "province": "",
  "teleworking": "Teletrabajo",
  "contractType": "Indefinido",
  "workday": "Completa",
  "postedSince": 0,
  "titleExclude": [
    "becario",
    "practicas"
  ],
  "companyExclude": [],
  "extractSkills": true,
  "onlyNewSinceLastRun": false,
  "maxItems": 15,
  "timeoutSecs": 25,
  "cacheTtlSeconds": 1800
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("nomad-agent/infojobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("nomad-agent/infojobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call nomad-agent/infojobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=nomad-agent/infojobs-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/rQXhPiMSny2KtShmX/builds/b0uVlPNWhp9RWR51n/openapi.json
