# Hacker News Job Scraper: Who is Hiring Posts (`getascraper/hn-hiring-scraper`) Actor

Scrape Hacker News Who is Hiring job posts into structured JSON. Extract company, role, salary, remote status, tech stack, emails, and application URLs. Drop-in for Google Sheets, Airtable, and Zapier. Skip manual copy-paste.

- **URL**: https://apify.com/getascraper/hn-hiring-scraper.md
- **Developed by:** [GetAScraper](https://apify.com/getascraper) (community)
- **Categories:** Lead generation, AI, Social media
- **Stats:** 5 total users, 4 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.67 / 1,000 jobs

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 💻 HN Who is Hiring Scraper

Extract structured job postings from Hacker News monthly "Who is Hiring?" threads. Parse company, role, location, remote status, salary, technologies, emails, and application URLs from the largest organic tech job board on the internet.

Built on the official Hacker News Firebase API and Algolia Search API for reliable, rate-limit-free access to job data.

### 💡 Why use it?

- **Structured Data**: Extracts company, role, location, remote status, salary, technologies, emails, and URLs from unstructured HN comments
- **Auto-Discovery**: Automatically finds the latest "Who is Hiring?" posts without needing specific URLs
- **Historical Data**: Scrape multiple months back for trend analysis
- **Tech-Focused**: Identifies technologies mentioned in each job post for filtering and analysis
- **Contact Extraction**: Automatically finds email addresses and application URLs

### 🚀 How to use

1. Open the Actor in Apify Console.
2. Leave `startUrls` empty to auto-discover the latest hiring post, or provide specific HN post URLs.
3. Set `monthsBack` to scrape multiple months (max 12).
4. Set `maxJobsPerMonth` to limit results (0 = unlimited).
5. Optionally enable `includeReplies` to capture nested discussion threads.
6. Run the Actor and consume the output via Apify API, CSV, or JSON.

### 📋 Input fields

- `startUrls` (array, optional): Specific HN "Who is Hiring?" post URLs. If empty, auto-discovers the latest posts.
- `monthsBack` (integer): How many months of hiring posts to scrape when auto-discovering. Default: 1, Max: 12.
- `maxJobsPerMonth` (integer): Maximum job postings to extract per month. Default: 0 (unlimited).
- `includeReplies` (boolean): Whether to include nested replies/discussion threads. Default: false.
- `proxyConfiguration` (object): Proxy configuration for API requests. Optional - HN APIs are generally open.

### 📦 Output schema

Each dataset item represents one job posting:

```json
{
  "commentId": 22666455,
  "hnUser": "kfx",
  "postedAt": 1584984975,
  "postedAtIso": "2020-03-23T17:36:15.000Z",
  "rawText": "PBS | Various Engineers | Full-Time | ONSITE...",
  "cleanText": "PBS | Various Engineers | Full-Time | ONSITE...",
  "company": "PBS",
  "role": "Various Engineers",
  "location": "Alexandria, VA",
  "remoteStatus": "ONSITE (Flexible WFH)",
  "employmentType": "Full-Time",
  "salary": null,
  "technologies": ["express", "iOS"],
  "emails": ["digitaljobs@pbs.org"],
  "urls": ["https://tinyurl.com/v7c8nb2"],
  "isTopLevel": true,
  "parentId": 22665398,
  "replyCount": 0,
  "hnUrl": "https://news.ycombinator.com/item?id=22666455"
}
```

### 📊 Data table

| Field | Type | Description |
|---|---|---|
| `commentId` | number | Unique HN comment ID |
| `company` | string | Company name (extracted from post) |
| `role` | string | Job role/title |
| `location` | string | Job location |
| `remoteStatus` | string | REMOTE, ONSITE, HYBRID, etc. |
| `employmentType` | string | Full-time, Contract, Intern, etc. |
| `salary` | string | Salary range if found in text |
| `technologies` | array | Technologies mentioned in the post |
| `emails` | array | Email addresses found |
| `urls` | array | Application/company URLs found |
| `hnUser` | string | HN username who posted the job |
| `postedAtIso` | string | ISO timestamp of the post |
| `replyCount` | number | Number of replies to this job post |
| `hnUrl` | string | Direct link to the comment on HN |

### 💰 Pricing / cost estimation

Priced at **$0.02 per job posting** (Pay-per-Result).

| Target Jobs | Estimated Cost |
|---|---|
| 100 | $2.00 |
| 500 | $10.00 |
| 1,000 | $20.00 |

HN APIs are open and free to access. No proxy costs typically required.

### ⭐ Enjoying HN Who is Hiring Scraper?

<table width="100%">
<tr>
<td style="padding:20px 24px 14px;background:#FFF6F0;border:1px solid #FFF6F0;border-left:5px solid #FF6600;border-radius:10px 10px 0 0">
<span style="font-size:20px;letter-spacing:4px">⭐ ⭐ ⭐ ⭐ ⭐</span><br>
<span style="font-size:17px;font-weight:800;color:#1C1917">Turn Hacker News hiring threads into a searchable feed of company, role, and salary leads within minutes.</span><br>
<span style="font-size:14px;color:#57534E">A 5-star rating takes 10 seconds and helps other tech recruiters and job seekers find it. Your feedback also tells us what to build next.</span>
</td>
</tr>
<tr>
<td style="padding:0;background:#FF6600;border:1px solid #FFF6F0;border-top:none;border-radius:0 0 10px 10px;text-align:center">
<a href="https://apify.com/getascraper/hn-hiring-scraper/reviews" style="display:block;padding:13px 16px;color:#FFFFFF;text-decoration:none;font-weight:800;font-size:15px;letter-spacing:0.3px">★&nbsp;&nbsp;Rate this Actor on Apify</a>
</td>
</tr>
</table>

### ✨ Tips / Advanced

- **Auto-Discovery**: Leave `startUrls` empty and set `monthsBack` to 3-6 to get a rolling window of hiring posts
- **Focus on Remote**: Filter output by `remoteStatus` field containing "REMOTE"
- **Tech Filtering**: Use the `technologies` array to find jobs matching specific skills
- **Speed**: Each API call has a 50ms delay to be respectful to HN. Expect ~20 jobs/minute

### ❓ FAQ

**Is scraping Hacker News legal?**
HN provides official APIs (Firebase and Algolia) for accessing this data. This Actor uses those APIs, not HTML scraping.

**Why did I get fewer results than expected?**
Some comments in hiring threads are discussion, not job posts. The parser attempts to filter these, but imperfectly. Set `includeReplies: true` to capture more.

**Can I scrape historical data?**
Yes. Set `monthsBack` up to 12 to scrape past hiring threads. Note that older posts may have fewer active listings.

### 🛠️ Support

For bug reports or feature requests, open a ticket in the Issues tab.

### 🔗 Other actors

- [TeamBlind Reviews Scraper](https://apify.com/getascraper/teamblind-reviews-scraper) ↗ - collects anonymous tech employer ratings and reviews from Blind.
- [NoFluffJobs Scraper](https://apify.com/getascraper/nofluffjobs-scraper) ↗ - extracts tech job listings and transparent salary ranges.
- [CWJobs Scraper](https://apify.com/getascraper/cwjobs-scraper) ↗ - pulls UK tech job listings with salaries and locations.
- [Wantedly Japan Jobs Scraper](https://apify.com/getascraper/wantedly-jobs-scraper) ↗ - gathers startup job listings from Japan's Wantedly platform.
- [Bug Bounty Finder](https://apify.com/getascraper/bug-bounty-finder) ↗ - aggregates active bug bounty programs from HackerOne and Bugcrowd.

# Actor input Schema

## `startUrls` (type: `array`):

Specific Hacker News 'Who is Hiring?' post URLs. If empty, auto-discovers the latest posts.

## `monthsBack` (type: `integer`):

How many months of hiring posts to scrape when auto-discovering. Max 12.

## `maxJobsPerMonth` (type: `integer`):

Maximum job postings to extract per month. 0 = unlimited.

## `includeReplies` (type: `boolean`):

Whether to include nested replies/discussion threads.

## `proxyConfiguration` (type: `object`):

Optional proxy settings. Defaults to enabling Apify Proxy to prevent blocks during cloud runs.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://news.ycombinator.com/item?id=47975571"
    }
  ],
  "monthsBack": 1,
  "maxJobsPerMonth": 0,
  "includeReplies": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://news.ycombinator.com/item?id=47975571"
        }
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("getascraper/hn-hiring-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://news.ycombinator.com/item?id=47975571" }],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("getascraper/hn-hiring-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://news.ycombinator.com/item?id=47975571"
    }
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call getascraper/hn-hiring-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=getascraper/hn-hiring-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/gYsH3egwBgpHrYvDy/builds/cSFLM5VaSNIQzoZmf/openapi.json
