# Technical SEO & AI Crawler Audit (`taroyamada/technical-seo-portfolio-regression-report`) Actor

Audit public pages you are authorized to test for robots, noindex, sitemap, canonical, and AI crawler access changes. Deliver source-linked issues and portfolio reports.

- **URL**: https://apify.com/taroyamada/technical-seo-portfolio-regression-report.md
- **Developed by:** [naoki anzai](https://apify.com/taroyamada) (community)
- **Categories:** Developer tools, Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $12.00 / 1,000 seo page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Technical SEO & AI Crawler Audit

The default `initialRunMode: "baseline_only"` stores the first portfolio baseline with zero rows and zero charge. Later unchanged monitor runs also emit zero rows and zero charge. Use `emit_backfill` only when you explicitly want the current observations and report delivered as billable rows.

Site owners, SEO agencies, and content teams use this actor to audit public indexability and AI crawler access signals.
Provide public URLs and optional AI crawler user-agent names.
The actor returns source-linked policy observations, indexability issues, reports, and export rows.

### Start with a low-cost paid trial

Use this capped input to verify the source and buyer-facing output before buying a full report. It disables the report and export events and caps the run at **$0.00**. This authorization-gated starter is a free safe demo; confirm authorization and choose a paid output before a live audit. A zero-row run has zero event charge.

```json
{
  "urls": [
    "https://example.com/"
  ],
  "aiCrawlerUserAgents": [
    "GPTBot",
    "ClaudeBot"
  ],
  "maxPages": 3,
  "checkRobotsTxt": true,
  "checkLlmsTxt": true,
  "authorizedUseConfirmed": false,
  "emitPageRows": false,
  "emitRawRows": false,
  "generateReport": true,
  "emitExport": false,
  "emitUnchanged": false,
  "initialRunMode": "baseline_only",
  "maxChargeUsd": 0,
  "dryRun": true
}
```

When the result fits your workflow, use the full-report input below to generate the full `seo-portfolio-report` output (**$3.00**) and `seo-agency-export` export (**$2.00**). The cap covers the selected report/export path; unchanged monitoring runs still produce zero rows and zero event charge.

```json
{
  "urls": [
    "https://your-authorized-domain.example/"
  ],
  "aiCrawlerUserAgents": [
    "GPTBot",
    "ClaudeBot"
  ],
  "maxPages": 3,
  "checkRobotsTxt": true,
  "checkLlmsTxt": true,
  "authorizedUseConfirmed": false,
  "emitPageRows": false,
  "emitRawRows": false,
  "generateReport": true,
  "emitExport": true,
  "emitUnchanged": false,
  "initialRunMode": "emit_backfill",
  "maxChargeUsd": 5.27,
  "dryRun": false
}
```

### Store Quickstart

Run with `dryRun=false` and public URLs that you own or are allowed to audit.

```json
{
  "urls": ["https://example.com/?siteQaCanary=indexability-ai-crawler-v1"],
  "aiCrawlerUserAgents": ["GPTBot", "Google-Extended", "PerplexityBot", "ClaudeBot"],
  "checkRobotsTxt": true,
  "checkLlmsTxt": true,
  "authorizedUseConfirmed": true,
  "generateReport": true,
  "emitUnchanged": false,
  "dryRun": false
}
```

### Run the next report

- `seo-portfolio-report`: $3.000 per portfolio report.
- `seo-agency-export`: $2.000 per generated export.
- Measure Core Web Vitals, Lighthouse scores, axe-core findings, and score regressions with [Website Lighthouse & WCAG Regression Report](https://apify.com/taroyamada/website-lighthouse-accessibility-regression-report).

### Input Examples

#### Audit one page and origin policies

```json
{
  "urls": ["https://example.com/blog/launch"],
  "aiCrawlerUserAgents": ["GPTBot", "ClaudeBot"],
  "checkRobotsTxt": true,
  "checkLlmsTxt": true,
  "authorizedUseConfirmed": true,
  "dryRun": false
}
```

#### Batch audit a site section

```json
{
  "urls": [
    "https://example.com/",
    "https://example.com/pricing",
    "https://example.com/docs"
  ],
  "maxPages": 25,
  "emitPageRows": false,
  "generateReport": true,
  "authorizedUseConfirmed": true,
  "dryRun": false
}
```

#### Generate a handoff export

```json
{
  "urls": ["https://example.com/landing-page"],
  "aiCrawlerUserAgents": ["GPTBot", "Google-Extended", "PerplexityBot"],
  "emitExport": true,
  "emitUnchanged": false,
  "authorizedUseConfirmed": true,
  "dryRun": false
}
```

### Sample Output

```json
{
  "actorName": "technical-seo-portfolio-regression-report",
  "rowType": "indexability_issue",
  "billingEventName": "seo-regression-detected",
  "issueType": "ai_crawler_disallowed_by_robots",
  "severity": "high",
  "sourceUrl": "https://example.com/?siteQaCanary=indexability-ai-crawler-v1"
}
```

### Output Fields

- `rowType`: `ai_crawler_policy_observation`, `indexability_issue`, `ai_crawler_indexability_report`, or `indexability_export`.
- `billingEventName`: PAY\_PER\_EVENT event name used for the row.
- `sourceUrl`: public URL or policy file that supports the row.
- `issueType`: detected source-linked issue when applicable.
- `blockedUserAgents`: crawler names with broad robots.txt blocks when detected.

### Pricing And No-Change Runs

- `seo-page-audited`: $0.012 per audited public page.
- `seo-regression-detected`: $0.080 per source-linked regression.
- `seo-portfolio-report`: $3.000 per portfolio report.
- `seo-agency-export`: $2.000 per generated export.

When `emitUnchanged=false`, repeated unchanged runs emit zero dataset rows and zero charges after state is saved.

`maxChargeUsd` is an optional positive per-run cap. The actor estimates every output row from the configured PAY\_PER\_EVENT unit prices before calling `pushData`; if the estimate exceeds the cap, it fails before charging. `0` preserves the existing no-explicit-cap behavior. Delivery journal entries record pending, charged, failed, and uncertain attempts so a partial retry can deduplicate confirmed charges and refuses to retry an uncertain charge.

### Compliance Guardrails

- Public pages, robots.txt, and llms.txt only.
- No login, paywall, CAPTCHA, private session, credentialed API, or bypass behavior.
- Non-dry runs require `authorizedUseConfirmed=true`; use this only for sites you own, manage, or are allowed to audit.
- This is an unofficial audit tool and is not affiliated with any crawler, search engine, or AI provider.
- No ranking guarantee, AI citation guarantee, legal conclusion, or compliance certification.

### Bundle Paths

- `seo-portfolio-report`: $3.000 per portfolio report.
- Pair with [Site QA Content Report Scraper](https://apify.com/taroyamada/site-qa-content-report-scraper) for page content issue reports.
- Pair with [Site QA Broken Link Report Scraper](https://apify.com/taroyamada/site-qa-broken-link-report-scraper) for link health reports.

### See Also

- [Site QA Content Report Scraper](https://apify.com/taroyamada/site-qa-content-report-scraper) for content QA issue reports.
- [Site QA Broken Link Report Scraper](https://apify.com/taroyamada/site-qa-broken-link-report-scraper) for broken link reports.

# Actor input Schema

## `urls` (type: `array`):

Public URLs that you are allowed to audit.

## `aiCrawlerUserAgents` (type: `array`):

Crawler user-agent names to inspect in robots.txt.

## `maxPages` (type: `integer`):

Maximum number of input pages to check.

## `checkRobotsTxt` (type: `boolean`):

Fetch and inspect robots.txt at each site origin.

## `checkLlmsTxt` (type: `boolean`):

Fetch and inspect llms.txt at each site origin.

## `authorizedUseConfirmed` (type: `boolean`):

Required for non-dry runs. Confirms each URL is owned by you, your client, or otherwise authorized for this audit.

## `emitPageRows` (type: `boolean`):

Emit optional public page indexability snapshot rows.

## `generateReport` (type: `boolean`):

Generate a site-level SEO portfolio report row.

## `emitExport` (type: `boolean`):

Generate an export row for handoff workflows.

## `emitUnchanged` (type: `boolean`):

When false, repeated unchanged runs emit zero rows and zero charges.

## `dryRun` (type: `boolean`):

Emit local sample rows without charging.

## `initialRunMode` (type: `string`):

Initial run mode for this run.

## `snapshotKey` (type: `string`):

Optional state namespace for canary and recurring no-change proof runs.

## `watchlists` (type: `array`):

Named portfolio watch definitions.

## `monitorKey` (type: `string`):

Monitor key for this run.

## `emitRawRows` (type: `boolean`):

Emit optional public page indexability snapshot rows.

## `maxChargeUsd` (type: `number`):

Maximum charge (USD) for this run.

## Actor input object example

```json
{
  "urls": [
    "https://example.com/"
  ],
  "aiCrawlerUserAgents": [
    "GPTBot",
    "ClaudeBot"
  ],
  "maxPages": 3,
  "checkRobotsTxt": true,
  "checkLlmsTxt": true,
  "authorizedUseConfirmed": false,
  "emitPageRows": false,
  "generateReport": true,
  "emitExport": false,
  "emitUnchanged": false,
  "dryRun": true,
  "initialRunMode": "baseline_only",
  "snapshotKey": "",
  "watchlists": [],
  "monitorKey": "",
  "emitRawRows": false,
  "maxChargeUsd": 0
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("taroyamada/technical-seo-portfolio-regression-report").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("taroyamada/technical-seo-portfolio-regression-report").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call taroyamada/technical-seo-portfolio-regression-report --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=taroyamada/technical-seo-portfolio-regression-report",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/eIARUlhXP8Se09G2P/builds/8XdUbfYMakKaMKk25/openapi.json
