# Career Site Job Listing Scraper (`thirdwatch/career-site-job-scraper`) Actor

Scrape job listings from any company career page. Auto-detects Lever, Greenhouse, Workday, BambooHR, Keka, and custom career sites. Extracts title, location, department, and apply URL.

- **URL**: https://apify.com/thirdwatch/career-site-job-scraper.md
- **Developed by:** [Thirdwatch](https://apify.com/thirdwatch) (community)
- **Categories:** Jobs
- **Stats:** 4 total users, 2 monthly users, 80.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.80 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Career Site Job Listing Scraper

> Track buying signals from company hiring: pull live job postings from Lever, Greenhouse, Workable, and 30+ ATS platforms. The cleanest "they're hiring for X" signal in B2B sales.

### Track hiring intent across 30+ ATS platforms

Lever, Greenhouse, Workday, BambooHR, Keka, Ashby, Recruitee, and any generic career page — one actor, one schema. Career sites are the source of truth for hiring intent: postings appear here days before LinkedIn Jobs or Indeed pick them up, and senior/stealth roles often never leave the employer site at all.

### Buying signals from job postings (the smartest B2B intent data)

If a company is hiring 5 Salesforce admins, sell them Salesforce. If they posted a "Senior Snowflake Engineer", they're a Snowflake account. Job postings are the most specific, real-time intent signal available — far more actionable than generic ZoomInfo Intent or Bombora topic scores. Pipe career-site data into your CRM and trigger plays the same day a relevant role goes live.

### What you get

Works with all major career page platforms: Lever, Greenhouse, Workday, BambooHR, Keka, Ashby, Recruitee, and any generic career page. Extracts job titles, locations, departments, and apply URLs automatically. Pass a list of company career URLs and get a clean, structured feed back — no per-platform configuration needed.

### Use cases & recipes

Step-by-step guides on [thirdwatch.dev/blog](https://thirdwatch.dev/blog):

- [Build a Jobs Aggregator from Company Career Pages (2026)](https://thirdwatch.dev/blog/build-jobs-aggregator-from-company-career-pages)
- [Scrape Greenhouse Jobs for ATS Enrichment (2026 Guide)](https://thirdwatch.dev/blog/scrape-greenhouse-jobs-for-ats-enrichment)
- [Scrape Lever Jobs for a Recruiter Sourcing Pipeline (2026)](https://thirdwatch.dev/blog/scrape-lever-jobs-for-recruiter-pipeline)
- [Track Startup Hiring Velocity with Career Site Data (2026)](https://thirdwatch.dev/blog/track-startup-hiring-velocity-with-career-sites)

### Supported platforms

- **Lever** (`jobs.lever.co/{company}`)
- **Greenhouse** (`boards.greenhouse.io/{company}`)
- **Workday** (`{company}.myworkdayjobs.com`)
- **BambooHR** (`{company}.bamboohr.com/careers`)
- **Keka** (`{company}.keka.com/careers`)
- **Ashby** (`jobs.ashbyhq.com/{company}`)
- **Recruitee** (`{company}.recruitee.com`)
- **Generic** (any career page with job listings)

### Output fields

| Field | Description |
|-------|-------------|
| `title` | Job title |
| `company_name` | Company name (auto-detected) |
| `department` | Department / team |
| `location` | Job location |
| `job_type` | Full-time, part-time, contract, etc. |
| `apply_url` | Direct application URL |
| `description` | Job description (when `scrapeDescriptions` is on) |
| `ats_platform` | Detected platform (lever, greenhouse, workday, etc.) |

### Example output

```json
{
    "title": "Senior Backend Engineer",
    "company_name": "Stripe",
    "department": "Engineering",
    "location": "San Francisco, CA",
    "job_type": "Full-time",
    "apply_url": "https://jobs.lever.co/stripe/abc123",
    "ats_platform": "lever"
}
```

### Input parameters

| Parameter | Required | Description |
|-----------|----------|-------------|
| `careerPageUrls` | Yes | Career page URLs to scrape (e.g., `https://jobs.lever.co/stripe`, `https://boards.greenhouse.io/discord`). |
| `companyName` | No | Override the auto-detected company name. Leave blank to auto-detect. |
| `scrapeDescriptions` | No | Visit each job detail page to extract full descriptions. Slower. Default `false`. |
| `maxJobsPerSite` | No | Maximum job listings per career page URL. Default `10`, max `500`. |
| `proxyConfiguration` | No | Pre-configured for best results. Leave as-is. |

### Use cases

- **Sales prospecting on hiring intent**: They're hiring 5 Salesforce admins — sell them Salesforce. Trigger outbound the day a relevant role posts.
- **Recruiter sourcing**: Find what companies are hiring for in real-time, before LinkedIn Jobs or Indeed mirror the postings.
- **Investor / equity research**: Track headcount growth across a portfolio as a leading indicator of revenue and burn.
- **Competitive intelligence**: See what your competitors are building from the engineering, product, and GTM roles they post.
- **Job aggregators**: Build a multi-company job feed sourced directly from employer career pages.
- **HR and talent analysts**: Track open positions at target companies and teams over time.
- **AI agents**: Feed an agent live "who's hiring for X" data straight from the source.

### Limitations

- Generic career pages (not on a supported platform) may return partial fields.
- Description scraping requires visiting each job page individually and adds runtime.
- Some Workday tenants use non-standard authentication and may not return data without additional configuration.
- Location strings vary by employer (city vs. city + country vs. "Remote") — no normalization.
- The actor reads public career listings; authenticated or internal-only jobs are not accessible.

### Compared to alternatives

- **vs. per-platform scrapers (Lever-only, Greenhouse-only, etc.)**: One actor covers 7+ platforms automatically — no need to glue together multiple scrapers.
- **vs. LinkedIn Jobs**: LinkedIn lags employer career pages by days and misses ATS-direct postings entirely (especially senior/stealth roles). This actor goes straight to the source so you catch roles the moment they go live.
- **vs. paid hiring-signal vendors (ZoomInfo Intent, Bombora, Demandbase)**: Those products charge $20K-$100K/yr for derived "intent" scores. Career-site data is the underlying primary signal — cheaper, raw, and far more specific (job title and department, not a topic bucket).

Pairs well with [LinkedIn Jobs Scraper](https://apify.com/thirdwatch/linkedin-jobs-scraper?fpr=9m2cd6) and [Indeed Jobs Scraper](https://apify.com/thirdwatch/indeed-jobs-scraper?fpr=9m2cd6) for maximum-coverage job feeds.

### FAQ

**Do I need to know which platform each company uses?**
No. The scraper auto-detects Lever, Greenhouse, Workday, BambooHR, Keka, Ashby, Recruitee, and generic pages from the URL.

**Can I get full descriptions?**
Yes — set `scrapeDescriptions: true`. It fetches each job's detail page, so runs take longer.

**How do I monitor many companies?**
Schedule the actor daily or hourly on the Apify platform. Pass all career URLs in one run and downstream-dedupe by `apply_url`.

**How fresh is the data?**
Pulled live from the company's career page at run time.

Last verified: 2026-05

More scrapers at [thirdwatch.dev](https://thirdwatch.dev).

# Actor input Schema

## `sourceMode` (type: `string`):

Search the 2.5M+ direct-ATS index, or scrape specific career page URLs.

## `query` (type: `string`):

Role, skill, company, or free-text job search.

## `country` (type: `string`):

Country name or two-letter code, for example US, GB, Germany, or India.

## `city` (type: `string`):

Optional city filter, for example London, Bengaluru, or New York.

## `company` (type: `string`):

Optional employer name filter.

## `atsSource` (type: `string`):

Optional source filter such as workday, greenhouse, lever, ashby, or smartrecruiters.

## `postedSince` (type: `string`):

ISO date/time, for example 2026-07-01T00:00:00Z.

## `hasSalary` (type: `boolean`):

When enabled, exclude jobs without structured salary data.

## `minimumSalary` (type: `number`):

Optional minimum structured salary value.

## `employmentType` (type: `string`):

For example full\_time, part\_time, contract, or internship.

## `workArrangement` (type: `string`):

For example remote, hybrid, or on\_site.

## `excludeDuplicates` (type: `boolean`):

Collapse the same opening found through multiple source URLs.

## `includeDescriptions` (type: `boolean`):

Include full job descriptions instead of compact snippets.

## `maxResults` (type: `integer`):

Maximum number of matching jobs returned across all API pages.

## `careerPageUrls` (type: `array`):

List of company career page URLs to scrape (e.g., https://jobs.lever.co/stripe, https://boards.greenhouse.io/discord)

## `companyName` (type: `string`):

Override auto-detected company name. Leave blank to auto-detect from URL/page.

## `scrapeDescriptions` (type: `boolean`):

Visit each job detail page to extract full descriptions. Slower and uses more credits.

## `maxJobsPerSite` (type: `integer`):

Maximum number of job listings to extract per career page URL.

## `proxyConfiguration` (type: `object`):

Proxy settings. Leave default for best results.

## Actor input object example

```json
{
  "sourceMode": "careerPages",
  "query": "data engineer",
  "hasSalary": false,
  "excludeDuplicates": true,
  "includeDescriptions": false,
  "maxResults": 100,
  "careerPageUrls": [
    "https://jobs.lever.co/solopulseco",
    "https://boards.greenhouse.io/mark43"
  ],
  "scrapeDescriptions": false,
  "maxJobsPerSite": 10,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "data engineer",
    "maxResults": 100,
    "careerPageUrls": [
        "https://jobs.lever.co/solopulseco",
        "https://boards.greenhouse.io/mark43"
    ],
    "maxJobsPerSite": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("thirdwatch/career-site-job-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "data engineer",
    "maxResults": 100,
    "careerPageUrls": [
        "https://jobs.lever.co/solopulseco",
        "https://boards.greenhouse.io/mark43",
    ],
    "maxJobsPerSite": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("thirdwatch/career-site-job-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "data engineer",
  "maxResults": 100,
  "careerPageUrls": [
    "https://jobs.lever.co/solopulseco",
    "https://boards.greenhouse.io/mark43"
  ],
  "maxJobsPerSite": 10
}' |
apify call thirdwatch/career-site-job-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=thirdwatch/career-site-job-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/9iXGuFRD3yb5S0cAN/builds/acXfcAAWOInuEPjMC/openapi.json
