# Greenhouse Job Listings API – Careers Board Scraper (`deadwood_data_solutions/greenhouse-jobs-scraper`) Actor

Scrape live job postings from any company's Greenhouse board. Clean JSON per job with dedup, so scheduled runs only return newly-posted roles. For recruiters, ATS aggregators and hiring-signal feeds.

- **URL**: https://apify.com/deadwood\_data\_solutions/greenhouse-jobs-scraper.md
- **Developed by:** [K O](https://apify.com/deadwood_data_solutions) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Greenhouse Jobs Scraper – New Postings Feed API

Scrape live job postings from **any company's Greenhouse job board** and get one clean JSON record per role — title, location, departments, offices, requisition id, post/update dates and a direct apply URL. Built-in deduplication means a **scheduled run only returns newly-posted jobs**, so you can wire it straight into a hiring-signal feed, a niche job board, or a sourcing tool.

### Who uses it

Recruiters and sourcers, job-board and ATS aggregators, sales teams tracking a target account's hiring, labor-market analysts, and anyone building a "new roles at companies I follow" alert.

### Why this is worth charging for

Greenhouse exposes a public board API, but you still have to discover each company's board token, page through boards one at a time, decode HTML-entity job descriptions, and diff runs yourself to find what's actually new. This actor does all of that: pass a list of companies, get a normalized, deduplicated feed you can schedule.

### Input

| Field | Description |
|---|---|
| `companies` | One or more Greenhouse board tokens — the slug in `job-boards.greenhouse.io/<token>` (e.g. `airbnb`, `stripe`, `gitlab`). |
| `includeDescription` | Fetch each posting's full HTML description. Off by default for lean records. |
| `onlyNewSinceLastRun` | Recommended for schedules — skips jobs already returned by a previous run, so you're only charged for genuinely new postings. |
| `maxItems` | Stop after this many normalized records. |

### Output

Each dataset item is one normalized job:

```json
{
  "jobId": "7996480",
  "company": "airbnb",
  "title": "Acquisition Manager",
  "location": "Milan, Italy",
  "departments": ["Sales"],
  "offices": ["Milan"],
  "updatedAt": "2026-06-10T09:22:20-04:00",
  "firstPublished": "2026-01-02T00:00:00-05:00",
  "requisitionId": "R-123",
  "url": "https://careers.airbnb.com/positions/7996480",
  "descriptionHtml": null,
  "source": "Greenhouse"
}
```

### Pricing (Pay-Per-Event)

- **`query`** — charged once per run for the board poll.
- **`job-record`** — charged per normalized job pushed. This is the primary event.
- `apify-actor-start` (Apify-managed) — covers baseline compute per run.

A daily monitor of a handful of companies returns a few new roles for pennies; a full-board pull of a large employer is a larger one-time run you control with `maxItems`.

### Source & reliability

Data comes from Greenhouse's public `boards-api.greenhouse.io` service. No API key, no proxy needed. Run `npm test` for the offline self-test covering the normalizer and its edge cases.

### FAQ

**Where do I find a company's board token?**

Open the company's Greenhouse careers page — the token is the slug in the URL, e.g. `job-boards.greenhouse.io/airbnb` → `airbnb`.

**Can I get full job descriptions?**

Yes — set `includeDescription` to true and each record includes the decoded HTML description.

**How does it avoid charging me for the same job twice?**

With `onlyNewSinceLastRun` on (the default), the actor persists the ids it has already returned and skips them next run, so a schedule only bills for new postings.

***

*SEO keywords: Greenhouse jobs scraper, Greenhouse board API, job posting scraper, ATS scraper, new jobs feed, hiring signals, recruiting data API, job board aggregator*

# Actor input Schema

## `companies` (type: `array`):

One or more Greenhouse board tokens — the company slug in job-boards.greenhouse.io/<token> (e.g. airbnb, stripe, gitlab).

## `includeDescription` (type: `boolean`):

Fetch each posting's full HTML job description. Larger records and a slightly slower run.

## `maxItems` (type: `integer`):

Stop after this many normalized records have been pushed. Leave blank for no limit.

## `onlyNewSinceLastRun` (type: `boolean`):

Recommended for scheduled runs. Uses persisted state to skip postings already returned earlier, so a recurring schedule only charges for genuinely new jobs.

## Actor input object example

```json
{
  "companies": [
    "airbnb",
    "stripe"
  ],
  "includeDescription": false,
  "onlyNewSinceLastRun": true
}
```

# Actor output Schema

## `results` (type: `string`):

All normalized records from this run as JSON.

## `resultsCsv` (type: `string`):

All normalized records from this run as CSV.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "airbnb",
        "stripe"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("deadwood_data_solutions/greenhouse-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "companies": [
        "airbnb",
        "stripe",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("deadwood_data_solutions/greenhouse-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "airbnb",
    "stripe"
  ]
}' |
apify call deadwood_data_solutions/greenhouse-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=deadwood_data_solutions/greenhouse-jobs-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Gq52Ob9aI1NCCOrcG/builds/5gmXXf3pTf8aeh25X/openapi.json
