# LinkedIn Jobs Scraper (`mighty_monk/linkedin-jobs-scraper`) Actor

Scrape public LinkedIn job search results and job detail pages. Extract title, company, location, employment type, seniority, description, apply URL, posted date, and applicants info when public.

- **URL**: https://apify.com/mighty\_monk/linkedin-jobs-scraper.md
- **Developed by:** [Harsh](https://apify.com/mighty_monk) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 5 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 job scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## LinkedIn Jobs Scraper

**Scrape public LinkedIn job search results and job detail pages** — title, company, location, workplace type, employment type, seniority, description, apply URL, posted date, and applicants (when public). Built with **TypeScript**, **Crawlee CheerioCrawler**, and LinkedIn **guest/public** job endpoints (no login required for many public listings).

Run it on the [Apify platform](https://apify.com) for scheduling, API access, proxy rotation, monitoring, and dataset exports (JSON, CSV, Excel).

### What does LinkedIn Jobs Scraper do?

Given LinkedIn job **search URLs** and/or **keywords + location**, this Actor:

1. Converts public search URLs into LinkedIn guest `seeMoreJobPostings` API requests
2. Paginates search results until `maxJobs` is reached
3. Optionally opens each job’s guest detail page for full description and criteria
4. Deduplicates jobs by `jobId`
5. Pushes structured dataset items
6. Exports **JOBS.csv** and **SUMMARY.md** to the key-value store

### Why use LinkedIn Jobs Scraper?

- **Recruiting pipelines** — build candidate-role inventories from public listings
- **Sales / lead generation** — find companies hiring for target roles
- **Market research** — track job titles, locations, and seniority mix over time
- **Competitive intel** — monitor which skills and roles competitors post

### How to scrape LinkedIn jobs

1. Open the Actor in [Apify Console](https://console.apify.com)
2. Paste a LinkedIn jobs search URL **or** set `keywords` + `location`
3. Set **Max jobs** and enable/disable full descriptions
4. Keep concurrency low and enable **Apify Proxy** (residential recommended)
5. Click **Start**
6. Download results from the **Dataset** tab or key-value store exports

### Input

| Field                | Type           | Default        | Description                                      |
| -------------------- | -------------- | -------------- | ------------------------------------------------ |
| `searchUrls`         | string\[]       | `[]`           | Public LinkedIn job search URLs                  |
| `keywords`           | string         | `""`           | Keywords used to build a search URL              |
| `location`           | string         | `""`           | Location paired with keywords                    |
| `maxJobs`            | integer        | `50`           | Max unique jobs (by jobId)                       |
| `maxConcurrency`     | integer        | `2`            | Concurrent HTTP requests                         |
| `maxRequestRetries`  | integer        | `3`            | HTTP retry count                                 |
| `requestDelayMs`     | integer        | `1000`         | Delay used for rate limiting                     |
| `proxyConfiguration` | object         | Apify Proxy on | Proxy settings                                   |
| `includeDescription` | boolean        | `true`         | Fetch job detail pages for full description      |

Example (search URL):

```json
{
    "searchUrls": [
        "https://www.linkedin.com/jobs/search/?keywords=software%20engineer&location=United%20States"
    ],
    "maxJobs": 50,
    "maxConcurrency": 2,
    "maxRequestRetries": 3,
    "requestDelayMs": 1000,
    "proxyConfiguration": { "useApifyProxy": true },
    "includeDescription": true
}
```

Example (keywords + location):

```json
{
    "keywords": "account executive",
    "location": "New York, United States",
    "maxJobs": 25,
    "includeDescription": true,
    "proxyConfiguration": { "useApifyProxy": true }
}
```

See [`examples/input.json`](examples/input.json).

### Output

Each dataset item is one job:

```json
{
    "jobId": "4123456789",
    "title": "Senior Software Engineer",
    "company": "Acme Corp",
    "companyUrl": "https://www.linkedin.com/company/acme-corp",
    "location": "San Francisco, CA",
    "workplaceType": "Remote",
    "employmentType": "Full-time",
    "seniority": "Mid-Senior level",
    "description": "Build scalable backend services in TypeScript and Node.js...",
    "postedAt": "2026-07-01",
    "applicants": "Over 200 applicants",
    "applyUrl": "https://www.linkedin.com/jobs/view/4123456789/?isApplication=true",
    "jobUrl": "https://www.linkedin.com/jobs/view/4123456789",
    "scrapedAt": "2026-07-13T12:00:00.000Z",
    "error": null
}
```

You can download the dataset as **JSON, HTML, CSV, or Excel**. The Actor also writes:

| Key-value record | Description                        |
| ---------------- | ---------------------------------- |
| `JOBS.csv`       | All job rows as CSV                |
| `SUMMARY.md`     | Markdown summary table of results  |

### Data fields

| Field            | Description                                      |
| ---------------- | ------------------------------------------------ |
| `jobId`          | LinkedIn job posting ID                          |
| `title`          | Job title                                        |
| `company`        | Hiring company name                              |
| `companyUrl`     | LinkedIn company URL when available              |
| `location`       | Location string from the listing                 |
| `workplaceType`  | Remote / Hybrid / On-site (inferred when possible) |
| `employmentType` | Full-time, Part-time, Contract, etc.             |
| `seniority`      | Seniority level when public                      |
| `description`    | Full text when `includeDescription` is true      |
| `postedAt`       | Posted date or relative time                     |
| `applicants`     | Public applicants caption                        |
| `applyUrl`       | Application link when available                  |
| `jobUrl`         | Canonical public job URL                         |
| `scrapedAt`      | ISO scrape timestamp                             |
| `error`          | Soft error for blocked/partial records           |

### Pricing

**Pay-per-event (PPE):** **$0.003 per job** (dataset item), plus a small Actor-start fee.

Estimate: 50 jobs ≈ **$0.15**.

### Tips

- Prefer **Apify Proxy residential** if you hit auth walls, captchas, or empty results
- Keep `maxConcurrency` at **1–3** and `requestDelayMs` around **1000–2000 ms**
- Set `includeDescription: false` for faster/cheaper runs when you only need titles and companies
- Use stable public search URLs from an incognito browser session
- Unit tests use HTML fixtures so CI does not depend on live LinkedIn access

### Limitations

- Scrapes **public/guest** LinkedIn job pages only — not authenticated feed data
- LinkedIn frequently shows **auth walls** or rate limits; proxy + low concurrency help but do not guarantee 100% coverage
- Cheerio parses **static HTML**. Heavily client-rendered views may return fewer fields
- Some fields (applicants, apply URL, full description) are only present on detail pages
- Respect LinkedIn Terms of Service and applicable laws; use data responsibly

### Technical notes

- Search pagination uses\
  `https://www.linkedin.com/jobs-guest/jobs/api/seeMoreJobPostings/search?...&start=N`
- Job details use\
  `https://www.linkedin.com/jobs-guest/jobs/api/jobPosting/{jobId}`
- Extraction strategies: DOM job cards, JSON-LD `JobPosting`, job criteria lists
- Graceful handling of auth walls / blocks: logs a warning and falls back to search-card data when detail is blocked

### Development

```bash
cd factory/linkedin-jobs-scraper
npm install
npm test
npm run build
apify run
```

### License

Apache-2.0

# Actor input Schema

## `searchUrls` (type: `array`):

One or more public LinkedIn job search URLs (e.g. https://www.linkedin.com/jobs/search/?keywords=software%20engineer\&location=United%20States). The Actor converts these to guest API endpoints for scraping.

## `keywords` (type: `string`):

Job search keywords used when building a search URL (if searchUrls is empty or in addition to searchUrls). Example: software engineer.

## `location` (type: `string`):

Location string paired with keywords (e.g. United States, Remote, San Francisco, CA).

## `maxJobs` (type: `integer`):

Maximum number of unique jobs to scrape (deduplicated by jobId).

## `maxConcurrency` (type: `integer`):

Maximum number of concurrent HTTP requests. Keep low (1–3) to reduce LinkedIn rate limits and auth walls.

## `maxRequestRetries` (type: `integer`):

How many times to retry a failed HTTP request before giving up.

## `requestDelayMs` (type: `integer`):

Approximate delay between requests used to compute rate limiting (max requests per minute).

## `proxyConfiguration` (type: `object`):

Apify Proxy settings. Strongly recommended for LinkedIn (residential often works better).

## `includeDescription` (type: `boolean`):

When true, visit each job detail page (guest jobPosting API) to extract full description, seniority, employment type, applicants, and apply URL.

## Actor input object example

```json
{
  "searchUrls": [
    "https://www.linkedin.com/jobs/search/?keywords=software%20engineer&location=United%20States"
  ],
  "keywords": "",
  "location": "",
  "maxJobs": 50,
  "maxConcurrency": 2,
  "maxRequestRetries": 3,
  "requestDelayMs": 1000,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "includeDescription": true
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `jobsCsv` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchUrls": [
        "https://www.linkedin.com/jobs/search/?keywords=software%20engineer&location=United%20States"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mighty_monk/linkedin-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchUrls": ["https://www.linkedin.com/jobs/search/?keywords=software%20engineer&location=United%20States"] }

# Run the Actor and wait for it to finish
run = client.actor("mighty_monk/linkedin-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchUrls": [
    "https://www.linkedin.com/jobs/search/?keywords=software%20engineer&location=United%20States"
  ]
}' |
apify call mighty_monk/linkedin-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=mighty_monk/linkedin-jobs-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/15ojFVomb6iqq2l9e/builds/2k0YAXbWjEo9aSo1U/openapi.json
