# EURAXESS Jobs Scraper — EU Research Positions (`nomad-agent/euraxess-scraper`) Actor

Extract research vacancies from EURAXESS, the EU's researcher mobility portal: PhD, postdoc, fellowship and faculty roles across Europe. Each record carries title, institution, country, research field, contract type, deadline and apply URL. New-only delta alerts, no login required.

- **URL**: https://apify.com/nomad-agent/euraxess-scraper.md
- **Developed by:** [Nomad.Dev](https://apify.com/nomad-agent) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 16 total users, 6 monthly users, 64.7% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.10 / 1,000 job results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## EURAXESS Jobs Scraper — EU Research Positions

> **Claude / Codex skill to describe and setup this actor: [SKILL.md](https://github.com/Exdenta/OinkAIJobSearch/blob/main/skill/euraxess-scraper/SKILL.md)**

Fetch open research positions — PhD, postdoc, fellowship and faculty — from EURAXESS (also searched as "Euroaxess" or "Euraxxess"), the EU's official researcher mobility portal.

Why this scraper: up to 500 jobs per run, delta/alert mode that only bills genuinely new postings (ideal on a cron schedule), optional keyword translation across EU languages to catch local-language titles, and runs that never fail silently — an empty result always comes with an unbilled diagnostic row explaining why.

### What EURAXESS jobs data does this scraper extract?

Each result is one flat JSON record per job posting:

| Field | Meaning |
|---|---|
| `id` | Stable source-side identifier (EURAXESS node id) |
| `title` | Job title as posted |
| `company` | Hiring institution / organisation |
| `location` | City + country where stated, else country only |
| `country` | Posting country (EURAXESS highlight label) |
| `url` | Direct link to the posting |
| `postedAt` | Posting date (`YYYY-MM-DD`), `null` if not stated |
| `field` | Research discipline tag(s), e.g. `"Medical sciences » Health sciences"`; multiple tags joined with `; `, `null` if untagged |
| `contractType` | Contract type label (e.g. `"Temporary"`, `"Permanent"`, `"To be defined"`), `null` if the posting states none |
| `deadline` | Application deadline as an ISO 8601 timestamp, `null` if the posting has none |
| `snippet` | Short description excerpt |
| `isNew` | `true` when first seen on this run in delta mode (`onlyNewSinceLastRun`); absent otherwise |

### How to scrape EURAXESS jobs with this Actor

1. Click **Try for free** / **Run** — no login to the target site, no cookies, no proxies to configure.
2. Adjust the input (keyword, filters, `maxItems`) or keep the defaults.
3. Run it and export the dataset as JSON, CSV or Excel, or read it over the [API](https://docs.apify.com/api/v2).

Run it from your own code:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("nomad-agent/euraxess-scraper").call(run_input={"maxItems": 50})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["title"], "—", item["company"], item["url"])
```

Or a single HTTP call that runs the Actor and returns items in one response:

```bash
curl -X POST \
  "https://api.apify.com/v2/acts/nomad-agent~euraxess-scraper/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"maxItems": 50}'
```

### Input

| Field | Type | Default | Notes |
|---|---|---|---|
| `keyword` | string | `""` | Free-text search across job titles and descriptions (e.g. "machine learning", "postdoc biology"). Leave empty to return all current job offers. |
| `countryFilter` | string | `""` | Optional case-insensitive substring match on the posting country, applied client-side. E.g. "germany", "france", "spain". |
| `postedSince` | integer | `0` | Only return job offers posted within this many days (client-side, applied after fetch). `0` = no freshness filter. |
| `titleExclude` | array of strings | `[]` | Skip listings whose title contains any of these case-insensitive terms. |
| `companyExclude` | array of strings | `[]` | Skip listings whose company name contains any of these case-insensitive terms. |
| `maxItems` | integer | `100` | Maximum number of job offers to return. Each returned job is a billed result. `0` = no limit (the source exposes roughly 500 current offers per query). |
| `onlyNewSinceLastRun` | boolean | `false` | Delta / alert mode — see below. |
| `translateKeywords` | boolean | `false` | Expand the keyword into EU languages (BYOK) — see below. |
| `aiProvider` | string | `anthropic` | Which provider runs translation: `anthropic` or `mistral`. |
| `anthropicApiKey` / `mistralApiKey` | string | — | Your own API key (secret); only used when `translateKeywords` is on. |
| `aiModel` / `mistralModel` | string | Haiku / Small | Model used for translation for the chosen provider. |
| `requestTimeoutSecs` | integer | `30` | *(Advanced)* How many seconds to wait for the EURAXESS website to answer before giving up on that page. |
| `cacheTtlSeconds` | integer | `1800` | *(Advanced)* Reuse a page already fetched from EURAXESS for this many seconds instead of downloading it again on a quick re-run. `0` disables caching. |

### Delta mode / alerts (`onlyNewSinceLastRun`)

Turn on **Only new since last run** to make the Actor emit — and bill — only postings it hasn't seen on a previous run with the flag on. Already-seen postings are dropped before they're pushed, so a daily/weekly scheduled run pays only for genuinely new research openings, and each new row carries `"isNew": true`. State is tracked per Actor in a dedicated key-value store, keyed by each posting's stable EURAXESS node id — so re-titled or re-listed postings aren't re-charged.

### Keyword translation (`translateKeywords`, bring your own key)

EURAXESS lists positions from across the EU, and many carry **local-language titles** — a Spanish `Investigador postdoctoral`, a German `Postdoktorand` — that an English-only keyword search misses. Turn on **Translate keyword** and the Actor expands your keyword into its equivalents across the major EU research languages via the Anthropic or Mistral API, then matches any of them against each listing. It's off by default and needs your own API key (billed by that provider, not by this Actor); without a key the run falls back to a normal keyword search and says so in a diagnostic row.

### Output example

```json
{
  "id": "449831",
  "title": "University Professor",
  "company": "State University of Applied Sciences in Przemyśl",
  "location": "podkarpackie, Poland",
  "country": "Poland",
  "url": "https://euraxess.ec.europa.eu/jobs/449831",
  "postedAt": "2026-07-02",
  "field": "Medical sciences » Health sciences",
  "contractType": "Temporary",
  "deadline": "2026-08-18T21:05:49+00:00",
  "snippet": "State University of Applied Sciences in Przemyśl announces a competition for the position of University Professor at the Institute of Physiotherapy..."
}
```

### Integrations

Export the dataset as JSON, CSV or Excel from the Console, pull it over the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items` for a single-call run), wire it into Make/Zapier/n8n, or call it from AI agents via the [Apify MCP server](https://docs.apify.com/platform/integrations/mcp).

### Pricing

Pay per event: **$0.005 per Actor start** and **$0.004 per job returned**.
100 jobs ≈ $0.405. No subscription, no rental — you pay only for what you fetch.

### Use cases

- Academic job boards and PhD-alert bots
- University career services
- Research-mobility analytics across EU countries
- Grant and fellowship deadline tracking

### FAQ

**Is it legal to scrape EURAXESS jobs?**
This Actor reads only publicly available job postings — data any visitor can see without logging in. No personal data behind authentication is touched. Review the target site's terms and your local regulations for your specific use case.

**Do I need an account on the target site?**
No. Postings are fetched from public pages/APIs — no login, cookies or session tokens.

**How fresh is the data?**
Every run fetches live listings. Results are cached for `cacheTtlSeconds` (default 30 min, set 0 to always hit the source live).

**How many jobs can I get?**
`maxItems` caps the run (set 0 where supported for no cap). Most sources paginate from newest to oldest.

**Something broken or missing?**
Open an issue on the Actor's **Issues** tab — it is monitored and reliability fixes ship fast.

**Finding it useful?**
A short review on the Actor page genuinely helps other researchers and recruiters find it — and tells us which features to build next. Thank you!

### Related Actors

- [Research & Academic Jobs Scraper — 10 Sources](https://apify.com/nomad-agent/researcher-bundle)
- [jobs.ac.uk Scraper — UK Academic & Research Jobs](https://apify.com/nomad-agent/jobs-ac-uk-scraper)
- [Ikerbasque Jobs Scraper — Basque Research Roles](https://apify.com/nomad-agent/ikerbasque-scraper)
- [University of Copenhagen PhD Jobs Scraper (KU)](https://apify.com/nomad-agent/math-ku-phd-scraper)

***

**From the maker of [Oink](https://github.com/Exdenta/OinkAIJobSearch)** — an open-source, AI-powered job-search bot for Telegram that runs on these Actors. [Try the free bot](https://t.me/job_search_everyday_bot), get a managed instance at [oinkjobsearch.com](https://oinkjobsearch.com), or browse the [full catalog of 50+ Actors](https://apify.com/nomad-agent).

# Actor input Schema

## `keyword` (type: `string`):

Free-text search across job titles and descriptions (e.g. <code>machine learning</code>, <code>postdoc biology</code>). Leave empty to return all current job offers.

## `countryFilter` (type: `string`):

Optional case-insensitive substring match on the posting country, applied client-side. E.g. <code>germany</code>, <code>france</code>, <code>spain</code>.

## `postedSince` (type: `integer`):

Only return job offers posted within this many days (applied client-side after fetch — EURAXESS listings are already sorted newest first). Leave at 0 to return all current offers regardless of posting date.

## `titleExclude` (type: `array`):

Skip listings whose title contains any of these case-insensitive terms.

## `companyExclude` (type: `array`):

Skip listings whose company name contains any of these case-insensitive terms.

## `maxItems` (type: `integer`):

Maximum number of job offers to return. Each returned job is a billed result — see the Pricing tab. Set to 0 for no limit (the source exposes roughly 500 current offers per query).

## `onlyNewSinceLastRun` (type: `boolean`):

Delta / alert mode: only output postings not seen on a previous run that also had this flag on — the cheapest way to schedule this Actor on a cron and pay only for genuinely new research openings. Already-seen postings are dropped before push (not billed). State is tracked per Actor in a dedicated key-value store, keyed by each posting's EURAXESS node id. New rows carry <code>isNew: true</code>. See README "Delta mode / alerts".

## `translateKeywords` (type: `boolean`):

When on, expands your <b>Keyword</b> into its equivalents across the major EU research languages (German, French, Spanish, Italian…) via the Anthropic or Mistral API, then matches any of them against each listing — so a search for <code>postdoc</code> also catches local-language titles like <code>Postdoktorand</code> or <code>investigador postdoctoral</code>. Requires your own API key below (BYOK); if turned on without one, the run falls back to a normal keyword search and says so. Off by default.

## `aiProvider` (type: `string`):

Which AI provider runs keyword translation (only relevant when "Translate keyword" is on). <code>anthropic</code> (default) uses Claude via anthropicApiKey. <code>mistral</code> uses a Mistral model via mistralApiKey instead — pick this if you'd rather bring a Mistral key than an Anthropic one. <code>openai</code> uses a GPT model via openaiApiKey.

## `anthropicApiKey` (type: `string`):

Your Anthropic API key (sk-ant-…). Only used when "Translate keyword" is on and aiProvider is <code>anthropic</code>; billed separately by Anthropic. Not required unless translateKeywords is on.

## `aiModel` (type: `string`):

Claude model used for keyword translation when aiProvider is <code>anthropic</code> (only relevant when "Translate keyword" is on). Haiku is fast, cheap and ample for this short translation task; Sonnet gives higher-quality readings on unusual or multi-word phrases.

## `mistralApiKey` (type: `string`):

Your Mistral API key. Only used when "Translate keyword" is on and aiProvider is <code>mistral</code>; billed separately by Mistral. Not required unless translateKeywords is on with aiProvider=mistral.

## `mistralModel` (type: `string`):

Mistral model used for keyword translation when aiProvider is <code>mistral</code> (only relevant when "Translate keyword" is on). Small is the default — it matches larger Mistral models on this well-scoped translation task at a fraction of the cost.

## `openaiApiKey` (type: `string`):

Your OpenAI API key (sk-…). Only used when "Translate keyword" is on and aiProvider is <code>openai</code>; billed separately by OpenAI. Not required unless translateKeywords is on with aiProvider=openai.

## `openaiModel` (type: `string`):

OpenAI model used for keyword translation when aiProvider is <code>openai</code> (only relevant when "Translate keyword" is on). Defaults to <code>gpt-4.1-mini</code> — cheap, fast and ample for this short translation task. Any chat-completions model works, including the gpt-5 family (e.g. <code>gpt-5.4-mini</code>).

## `proxyConfiguration` (type: `object`):

Proxy used to reach EURAXESS. Leave at the default: the Actor starts on the fast datacenter proxy and switches itself to residential only if the portal blocks it, which is much quicker than forcing residential up front. Pick explicit proxy groups here and the Actor uses exactly those, with no automatic fallback.

## `requestTimeoutSecs` (type: `integer`):

How many seconds to wait for the EURAXESS website to answer before giving up on that page and moving on.

## `cacheTtlSeconds` (type: `integer`):

Reuse a page already fetched from EURAXESS for this many seconds instead of downloading it again on a quick re-run. Set to 0 to always fetch fresh.

## Actor input object example

```json
{
  "keyword": "data science",
  "countryFilter": "germany",
  "postedSince": 0,
  "titleExclude": [
    "internship",
    "traineeship"
  ],
  "companyExclude": [],
  "maxItems": 100,
  "onlyNewSinceLastRun": false,
  "translateKeywords": false,
  "aiProvider": "anthropic",
  "aiModel": "claude-haiku-4-5-20251001",
  "mistralModel": "mistral-small-latest",
  "openaiModel": "gpt-4.1-mini",
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "requestTimeoutSecs": 30,
  "cacheTtlSeconds": 1800
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("nomad-agent/euraxess-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "proxyConfiguration": { "useApifyProxy": True } }

# Run the Actor and wait for it to finish
run = client.actor("nomad-agent/euraxess-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call nomad-agent/euraxess-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=nomad-agent/euraxess-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/N3gbK5IYfJI62GZj3/builds/tsQxomLf5fmLYWBdN/openapi.json
