# Job Posting Signal Normalizer (`zentrafoundry/job-posting-signal-normalizer`) Actor

Normalize job datasets and detect hiring signals.

- **URL**: https://apify.com/zentrafoundry/job-posting-signal-normalizer.md
- **Developed by:** [Zentra](https://apify.com/zentrafoundry) (community)
- **Categories:** Automation, Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$10.00 / 1,000 result delivereds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Job Posting Signal Normalizer

Normalize Schema Org Jobposting, Greenhouse Job Board Api, Lever Postings Api and return hiring-signal rows with company, role, skills, seniority, remote type, and change state.

### What this Actor does

### What it does

- Processes configured public sources or user-provided records for focused job posting signal normalizer monitoring.
- Emits structured rows with source references, stable identifiers, confidence, warnings, and run summary fields.
- Supports sample-mode runs so Apify Store QA and first-time users can inspect output without depending on live third-party availability.

### What it does not do

- Does not scrape private, login-only, paywalled, or access-restricted data unless the user provides approved credentials for a source they control.
- Does not guarantee every field is available from every source; missing or blocked fields are returned as warnings or nulls.
- Does not make legal, financial, compliance, procurement, medical, safety, or regulatory decisions.

### Who this is for

Recruiting, sales intelligence, talent market research, RevOps, and workforce analytics teams use this actor when they need focused job posting signal normalizer output instead of a broad generic scraper or manual checking.

### Buyer outcomes

- Normalize job posting signal normalizer job data across scraper outputs, ATS feeds, or career pages.
- Prioritize hiring signals with company, role, location, remote type, seniority, skills, change state, confidence, and source URL.
- Route hiring-signal records into recruiting, sales intelligence, workforce research, CRM, or alerting workflows.

### Data sources

### Sources monitored

- [Schema Org Jobposting](https://schema.org/JobPosting)
- [Greenhouse Job Board Api](https://developers.greenhouse.io/job-board.html)
- [Lever Postings Api](https://github.com/lever/postings-api)

### Input

- `sourceMode`: use `sample` for a smoke run or `startUrls` for job pages, ATS feeds, career pages, or scraper-output URLs.
- `startUrls`: job posting, career page, ATS, company hiring, or job scraper dataset URLs.
- `sourceIds`: approved job, ATS, scraper, or career-page source identifiers.
- `maxItems`: bounded number of normalized job records or job changes to return.
- `sinceLastRun`: emit only new or changed job posting states when scheduled.
- `watchlistTerms`: company, role, skill, seniority, location, remote type, or department terms.
- `webhookUrl`: optional destination for recruiting, sales intelligence, or hiring-signal alerts.

### How it transforms the input

- Input: job scraper dataset, ATS feed, career page, or job posting URL.
- Transformation: normalize company, role, location, remote type, seniority, department, and skills, then detect posting changes.
- Output: hiring-signal record with company, job title, skills, seniority, remote type, change type, source URL, and confidence.

### Output

The actor returns normalized job posting signal records with company, job title, location, remote type, seniority, skills, change type, hiring signal, source URL, and confidence.

Family-specific fields to expect:

- `companyName`: Normalized hiring company name.

- `jobTitle`: Normalized job title.

- `location`: Location or market for the posting.

- `remoteType`: Remote, hybrid, onsite, or unknown state.

- `seniority`: Junior, mid, senior, executive, or unspecified seniority.

- `skills`: Detected skill keywords or role capabilities.

- `hiringSignal`: Signal classification, such as posting-observed, new-posting, or role-change.

- `changeType`: Observed job posting state change.

- `sourceUrl`: Source-backed job posting or dataset URL.

- `detectedAt`: Timestamp when the hiring signal was detected.

- `recordId`: Stable record ID for exports, dedupe, and downstream joins.

- `title`: Human-readable record title for review and export.

- `sourceName`: Source identifier used to trace where the record came from.

- `sourceUrl`: Direct source URL for review and audit.

- `dedupeKey`: Stable key used for delta mode and duplicate suppression.

- `retrievedAt`: Timestamp showing when the actor retrieved or generated this record.

- `score`: Normalized field for filtering, routing, or downstream review.

- `scoreReasons`: Buyer-readable explanation for the score or match.

- `confidence`: Normalized field for filtering, routing, or downstream review.

- `errors`: Normalized field for filtering, routing, or downstream review.

- `runSummary`: Run-level summary for counts, filters, charges, and next actions.

### Pricing

This actor uses Apify pay-per-event pricing. Current public listing guidance: $29-$49 / 1,000 launch validation records until public data proof is complete. Charges are tied to buyer-visible value events such as `job-normalized`, `job-change-detected`, `dataset-processed`, `record-saved`, `enriched-record`. Small validation runs are supported so you can inspect output before scaling a schedule.

- `job-normalized`: Charge after producing one normalized job row. Typical price: $0.005. A run that produces 10 matching records charges only for the matched buyer-value events and remains capped by the run limit.
- `job-change-detected`: Charge after producing one job posting change. Typical price: $0.010. A run that produces 10 matching records charges only for the matched buyer-value events and remains capped by the run limit.
- `dataset-processed`: Base charge when Job Posting Signal Normalizer writes a non-empty default dataset. Typical price: $0.011. A run that produces 10 matching records charges only for the matched buyer-value events and remains capped by the run limit.
- `record-saved`: Charge for each buyer-visible result saved by Job Posting Signal Normalizer. Typical price: $0.003. A run that produces 10 matching records charges only for the matched buyer-value events and remains capped by the run limit.
- `first-run-cap`: Recommended first run budget cap. Typical price: $2.000. Start with the default small run, inspect the dataset, then raise maxItems or schedule recurring runs.

### API example

```bash
curl -X POST "https://api.apify.com/v2/actors/zentrafoundry~job-posting-signal-normalizer/runs" \
+  -H "Authorization: Bearer $APIFY_TOKEN" \
+  -H "Content-Type: application/json" \
+  -d '{"maxItems":10,"sourceIds":["SCHEMA-ORG-JOBPOSTING","GREENHOUSE-JOB-BOARD-API","LEVER-POSTINGS-API"],"includeSourceUrls":true,"includeMatchReasons":true,"outputMode":"buyer-ready-records"}'
```

### Demo run

### Recommended first run

```json
{
    "maxItems": 10,
    "sourceIds": [
        "SCHEMA-ORG-JOBPOSTING",
        "GREENHOUSE-JOB-BOARD-API",
        "LEVER-POSTINGS-API"
    ],
    "includeSourceUrls": true,
    "includeMatchReasons": true,
    "outputMode": "buyer-ready-records"
}
```

### Sample output

Review the real stored sample from the latest successful quality run: https://zentra.nimblique.studio/external/actor-review/samples/job-posting-signal-normalizer.json

### Recommended public tasks

```json
[
    {
        "name": "Normalize 10 job postings",
        "description": "Low-cost validation run for checking company, role, skill, seniority, remote type, and source fields.",
        "input": {
            "maxItems": 10,
            "sourceIds": [
                "SCHEMA-ORG-JOBPOSTING",
                "GREENHOUSE-JOB-BOARD-API",
                "LEVER-POSTINGS-API"
            ],
            "includeSourceUrls": true,
            "includeMatchReasons": true,
            "outputMode": "buyer-ready-records",
            "actorSlug": "job-posting-signal-normalizer"
        }
    },
    {
        "name": "Daily hiring signal check",
        "description": "Recurring batch for job posting changes, role trends, and hiring signals.",
        "schedule": "Daily during local business hours",
        "input": {
            "maxItems": 25,
            "sourceIds": [
                "SCHEMA-ORG-JOBPOSTING",
                "GREENHOUSE-JOB-BOARD-API",
                "LEVER-POSTINGS-API"
            ],
            "includeSourceUrls": true,
            "includeMatchReasons": true,
            "outputMode": "buyer-ready-records",
            "actorSlug": "job-posting-signal-normalizer"
        }
    }
]
```

### Example use cases

- Normalize job posting signal normalizer job postings into consistent company, role, location, skill, and seniority fields.
- Send hiring-signal records into recruiting, sales intelligence, or market research workflows.
- Detect new, changed, or removed postings with source evidence and confidence.
- Schedule recurring checks for hiring momentum and role-family changes.

### Trust and compliance

- Uses Schema Org Jobposting, Greenhouse Job Board Api, Lever Postings Api.
- Keeps source URLs and source identifiers in output records for auditability.
- Does not require private credentials unless a source is explicitly configured for approved authenticated access.

### Reliability and QA

- Prefilled Apify Store QA input runs in sample mode and should finish within the automated quality window.
- Empty input is handled with deterministic sample or diagnostic output instead of a crash.
- Demo/sample runs suppress buyer-value charges while still writing representative dataset rows.
- Production runs use bounded `maxItems`, source references, warnings, and run summaries so blocked or changed targets are visible.

### Limitations

- Results depend on public-source availability, source uptime, and source update cadence.
- Public sources can revise records after publication; rerun scheduled tasks for fresh evidence.
- Scores and match reasons are decision-support signals, not legal, financial, procurement, medical, safety, or regulatory advice.
- Large production runs can cost more than the default smoke run; start small, inspect output, then scale schedules.

### Legal and responsible use

Use this Actor only for public data or data you are authorized to process. You are responsible for complying with applicable laws, marketplace terms, robots policies, privacy rules, and source-specific limits.

### Support

Open an issue on the Actor page with the run ID, input summary, expected result, and observed result. Do not include secrets, cookies, auth headers, or private account data.

### FAQ

**Can I run this without URLs?** Yes. The default `sample` mode is designed to succeed without user-supplied URLs, and URL-backed runs can use `startUrls` when needed.

**Can I schedule it?** Yes. Use `sinceLastRun`, `watchlistTerms`, and optional `webhookUrl` to turn the actor into a recurring alert or report workflow.

**How do I verify value before scaling?** Run the recommended first-run input, review the sample output fields, then increase `maxItems` or schedule recurring runs after the dataset matches your use case.

# Actor input Schema

## `searchKeywords` (type: `array`):

Buyer-relevant keywords used to focus the records returned by this product.

## `taskIntent` (type: `string`):

Stable product-specific purpose for this saved Apify task.

## `sourceMode` (type: `string`):

Sample emits public-safe Job Company Signal rows. Approved live source mode keeps the same output fields and only uses owner-approved public URLs.

## `outputMode` (type: `string`):

Use sample records for Apify Store QA or buyer-ready records for approved Job Company Signal delivery.

## `startUrls` (type: `array`):

Public URLs to use for live Job Company Signal extraction after source-policy approval. Leave empty for sample mode.

## `maxItems` (type: `integer`):

Caps the number of Job Company Signal rows written to the dataset.

## `perSourceLimit` (type: `integer`):

Caps validated rows from any one source before cross-source deduplication.

## `maxTotalChargeUsd` (type: `number`):

Buyer-selected spend ceiling; the Apify run-level maximum remains authoritative.

## `overallTimeoutSecs` (type: `integer`):

Stops additional source work once the bounded run deadline is reached.

## `requestTimeoutSecs` (type: `integer`):

Timeout applied independently to each approved source request.

## `maxRequestRetries` (type: `integer`):

Bounded retry count for transient source failures.

## `sinceLastRun` (type: `boolean`):

Uses stable Actor state to skip logical records delivered by earlier runs.

## `deltaMode` (type: `boolean`):

Preserves stable deduplication keys for recurring tasks and schedules.

## Actor input object example

```json
{
  "searchKeywords": [
    "source-backed signal",
    "buyer fit"
  ],
  "taskIntent": "buyer-ready-product-run",
  "sourceMode": "sample",
  "outputMode": "sample-records",
  "startUrls": [],
  "maxItems": 1,
  "perSourceLimit": 25,
  "maxTotalChargeUsd": 5,
  "overallTimeoutSecs": 900,
  "requestTimeoutSecs": 30,
  "maxRequestRetries": 2,
  "sinceLastRun": false,
  "deltaMode": true
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchKeywords": [
        "source-backed signal",
        "buyer fit"
    ],
    "taskIntent": "buyer-ready-product-run",
    "sourceMode": "sample",
    "outputMode": "sample-records",
    "startUrls": [],
    "maxItems": 1,
    "perSourceLimit": 25,
    "maxTotalChargeUsd": 5,
    "overallTimeoutSecs": 900,
    "requestTimeoutSecs": 30,
    "maxRequestRetries": 2,
    "sinceLastRun": false,
    "deltaMode": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("zentrafoundry/job-posting-signal-normalizer").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchKeywords": [
        "source-backed signal",
        "buyer fit",
    ],
    "taskIntent": "buyer-ready-product-run",
    "sourceMode": "sample",
    "outputMode": "sample-records",
    "startUrls": [],
    "maxItems": 1,
    "perSourceLimit": 25,
    "maxTotalChargeUsd": 5,
    "overallTimeoutSecs": 900,
    "requestTimeoutSecs": 30,
    "maxRequestRetries": 2,
    "sinceLastRun": False,
    "deltaMode": True,
}

# Run the Actor and wait for it to finish
run = client.actor("zentrafoundry/job-posting-signal-normalizer").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchKeywords": [
    "source-backed signal",
    "buyer fit"
  ],
  "taskIntent": "buyer-ready-product-run",
  "sourceMode": "sample",
  "outputMode": "sample-records",
  "startUrls": [],
  "maxItems": 1,
  "perSourceLimit": 25,
  "maxTotalChargeUsd": 5,
  "overallTimeoutSecs": 900,
  "requestTimeoutSecs": 30,
  "maxRequestRetries": 2,
  "sinceLastRun": false,
  "deltaMode": true
}' |
apify call zentrafoundry/job-posting-signal-normalizer --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=zentrafoundry/job-posting-signal-normalizer",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/sLOEHIHmNgFzaIrnz/builds/sa7tEud38c2Vyicxc/openapi.json
