Evidence-Quoted Page Classifier — SEO + Your Own Labels, No AI avatar

Evidence-Quoted Page Classifier — SEO + Your Own Labels, No AI

Pricing

from $7.00 / 1,000 pages

Go to Apify Store
Evidence-Quoted Page Classifier — SEO + Your Own Labels, No AI

Evidence-Quoted Page Classifier — SEO + Your Own Labels, No AI

Give it your own labels and a list of pages; it tells you which pages say which, quoting the exact sentence every time. It never guesses: no words on the page means it says nothing. Includes a marketing-and-SEO lens plus full on-page SEO signals. No AI — your content never leaves.

Pricing

from $7.00 / 1,000 pages

Rating

0.0

(0)

Developer

Noah Davidson

Noah Davidson

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Evidence-Quoted Page Classifier — read how a page persuades, quoted not guessed

Point it at any pages and it reads how each one sells you — urgency, scarcity, authority, social proof — with the exact sentence that proves every call, and nothing where it can't point to the words. Bring your own labels, or use the built-in marketing-and-SEO lens.

In plain words: you give it your own labels and a list of pages, and it tells you which pages say which — quoting the exact sentence that matched.

Two things it does:

It never guesses. If the words aren't on the page, it says nothing at all. Not a low confidence score — nothing. So a result is never something it thought was probably true.

When a passage sits between two of your labels, it names both instead of forcing one. "Limited time offer — only 3 spots left" reads as urgency and scarcity. Every other tool picks one and prints a confidence number beside it. This one reports urgency / scarcity and how many passages landed there, so you end a run knowing where your own labels overlap — usually the real finding about the writing (copy genuinely doing two jobs, or two labels that aren't distinct in your material), and exactly what picking a winner destroys.

Here is the second one in an actual result, because it is the part people have to see to believe. Every run writes this alongside your data:

"ambiguous_between": {
"authority / social_proof": 6, // 6 passages where the two readings split — expertise vs. the crowd
"risk_reversal / social_proof": 4,
"scarcity / urgency": 2
},
"ambiguous_total": 12 // against 15 classifications it placed cleanly

Twelve passages it refused to force, each one told to you as a pair — a map of where your own labels overlap in your own copy. You are not charged for a single one of them.

You can check the other claim yourself on any run: every quote is findable word for word in the page it names (whitespace normalized), and — with Learn from this run off, the default — nothing from the old set survives between runs, so there's no memory for anything to be invented out of.

Bring your own label set, or use the included marketing-and-SEO lens — every page on the internet is trying to move you (hurry, last chance, everyone's switching, experts agree), and that lens reads exactly how a given page does it, alongside full on-page SEO signals and, if you switch it on, multimodal capture — which renders the page in a real headless browser and reads text baked into images, so JavaScript-heavy pages and text baked into images are read when the render succeeds and the text is OCR-legible (a failed render or an unreadable image yields an empty vision_text, not an error). Marketers use it to study copy at scale; anyone can use it to see the machinery plainly. SEO is one application of the engine, not the engine.

No LLM, no model call: your content never leaves the container, the same pages always give the same answer, and there's no per-result AI cost, so it stays cheap across thousands of pages. (Multimodal capture uses a local OCR reader — a symbolic floor with optional local neural corroboration — still no network, still no third-party model.)


What it's used for

Use it when you want to:

You want to…Do this
Audit meta/heading/canonical/OG hygiene across a site or a competitor setseo mode on your URL list
See how competitors persuade — urgency vs. social proof vs. authority — with the exact copyseo mode; read marketing_framing
Quantify persuasion framing on landing pages, reproducibly, for CRO or content workseo mode; read marketing_framing_counts
Get clean page text, including from JavaScript-heavy pages or text baked into imagesscrape mode, multimodal: true
Tag any set of pages against your own labels and trigger phrasescode mode with lensJson
Audit a whole site rather than a hand-picked listpoint it at your sitemap, or switch on Follow links
Check what an LLM wrote — measure its framing deterministically, with the sentence quotedcode mode on the model's output
Prove to a client which rows came from your audit, unalteredany mode; set a lens key, check provenance

What it does not do:

  • A rank tracker, keyword-research, or backlink tool. It reads pages; it does not query search engines or third-party SEO indexes.
  • A tone or imagery reader. It classifies what is written. Framing carried purely by design, photography, or implication has no sentence to quote, so it is reported as nothing — on purpose.

Quick start (about 30 seconds)

  1. Press Start with the defaults. The input is prefilled with https://apify.com in seo mode — run it as-is to see the shape of a result before you commit a real list.
  2. Paste your own URLs into Pages to analyze. A typed list, an uploaded file, another Actor's output, or a link to your sitemap all work — sitemap indexes are expanded one level, so sitemap.xml gives you the site's page list in one field, up to your Max pages limit (default 50 — raise it for a full site).
  3. Pick the mode — leave it on SEO + marketing framing unless you specifically want raw text (Multimodal scrape) or your own label set (Classify against a lens).
  4. Run, then open the Dataset tab. One row per page in seo mode; one row per classification in code mode.
  5. Read marketing_framing first. Each entry names a persuasion move and quotes the sentence that earned it — checkable against the live page.
  6. Then open CODING_REPORT and read ambiguous_between. That's where your question wasn't sharp enough yet: authority / social_proof: 6 means six passages sat between those two labels rather than landing on one. Merge them, or split them, or accept both — then run again. That is the whole refinement loop, and the tool never makes that call for you.

Everything else — rendering, crawling, per-page timeout, page cap, custom lenses — sits under Advanced options and has a working default.


Using it on AI output

The properties this engine was built for — same answer every time, never guesses, quotes its evidence — turn out to matter most somewhere other than web pages: checking what a language model just wrote.

Point code mode at model output instead of URLs and you get the same contract. If a model writes "Limited-time offer — only a few seats left", you get back urgency and scarcity, each with the exact words that produced them, and you get the same answer next month.

The reason that matters is short:

You cannot measure drift with an instrument that drifts.

Using one language model to grade another gives you a judge that samples, shifts between versions, and can't tell you why it scored something. Your measurements move because the ruler moved, and you can't tell which. This engine has no model in it, so it cannot drift, and every score it gives you comes attached to the sentence that caused it. Run it across model versions, prompt revisions, or a fine-tune, and any change you see is a change in the output — not in the ruler.

The honest boundary. It measures how something is framed, not whether it is true. It will tell you that a sentence uses urgency, and quote it. It will not tell you the claim is correct. Truth-checking against sources is a different job, and anything promising both from one deterministic pass is overselling.


Covering a whole site

Two ways, and the first is usually the right one.

1. Point it at your sitemap. Put https://yoursite.com/sitemap.xml into Pages to analyze using the field's "load URLs from" option. A sitemap index (a sitemap of sitemaps) is expanded one level, so most sites resolve to their full page list from that one entry — up to your Max pages limit (default 50; raise it for a full site). You get exactly the pages the site publishes, in one pass, with no crawl to supervise.

2. Switch on Follow links. Under Advanced options, turn on Follow links and set Crawl depth (1 = your starting pages plus the pages they link to). Useful when there's no sitemap, or the sitemap is stale.

Crawling holds to four rules, so a run can't wander:

  • Same site only. Links to other hosts are never followed. (www. and a port are the same site; a subdomain is not.)
  • Max pages is a hard stop, and the queue itself is bounded by the remaining budget — the run can never line up more work than it's allowed to do.
  • robots.txt is honored for every link the crawler discovers on its own. URLs you supplied are fetched as asked — that's your call to make; the ones it found itself are not.
  • Nothing is visited twice, so a site's navigation can't loop it.

Your bill is unchanged in shape: one charge per page fetched, whether you named it or the crawler found it. Max pages is your spend cap either way.


What comes back

seo mode — one row per page

Search-engine signals: title, title_length, meta_description, meta_keywords, meta_robots, canonical, og (Open Graph), twitter (card tags), headings (the full h1–h6 tree), h1_count, word_count.

Content & link health: images_total, images_with_alt, images_alt_coverage (the fraction carrying alt text — your accessibility and image-SEO gap in one number), links_total, links_internal, links_external, top_keywords (term + count).

Marketing framing — the part no standard SEO tool gives you:

  • marketing_framing — a list of {dimension, label, evidence_quote, review_needed}. The quote is verbatim from the page, so every call is traceable to the copy that produced it.
  • marketing_framing_counts — the same result tallied per label, for sorting a competitor set at a glance.

On every row: provenance — a tamper-evident tag (see Provenance). And error if a page couldn't be fetched, so a failure is a visible row rather than a silent gap.

The six framing labels

The built-in lens (marketing_seo_v1) reads one dimension, persuasion_framing, with six labels. Trigger phrases below are a representative sample — the full marker list ships in the lens file, and the markers gather candidate sentences rather than decide them (see How a label is decided).

LabelThe move it namesTriggers include
urgencyAct before a clock runs outtoday only, limited time, act now, last chance, ends soon, expires
scarcitySupply is short, others are taking itlimited stock, few left, almost gone, selling fast, exclusive, running out
social_proofOther people already chose thistrusted by, thousands of, join millions, best seller, 5-star, as seen on
authorityAn expert or institution vouches for itcertified, award-winning, industry-leading, backed by research, endorsed
benefit_ledWhat you gain, framed around youyou get, save, boost, unlock, transform, effortless, improve your
risk_reversalDownside removed so the choice feels freemoney-back, risk-free, free trial, cancel anytime, no credit card

code mode — one row per classification

dimension, label, evidence_quote, source_url, source_title, review_needed, provenance. A page that yields nothing quotable emits a single summary row with classifications: 0 — you get a receipt for every page read, never a silent absence.

scrape mode — one row per page

url, final_url (after redirects), status, title, text, error. With multimodal: true you also get rendered, vision_text (text read off the rendered view), and screenshot_key.

Learning from your own runs (optional)

Switch on Learn from this run and the classifications a run confirms are kept in a key-value store on your own account, then read back at the start of your next run. Repeated runs over your own material stand on more ground each time.

What that file contains is worth being precise about, because it's your data: only the labels and verbatim quotes that are already in your dataset — nothing else, and nothing about how the engine decided. You could assemble the same file yourself with a short script over your own results; this just saves you the script. Nothing is retained by us, the store is yours to delete, and the setting is off by default.

Each run writes a LEARNING_REPORT receipt: how much it formed, how much of that was new, how much it carried forward, and how much it holds now. The examples themselves land in LEARNED_EXAMPLES. Nothing is ever evicted — the set only grows.

It is automatic, and it lives on your account. Switch on Learn from this run and the examples are written to a key-value store under your own Apify account at the end of the run, and read back at the start of the next one. There are no files to move and nothing to wire up. You can open it, export it, or delete it at any time; we keep no copy.

Each example remembers the question and the labels it was formed under. A different question, or a different label set, keeps its own separate ground in the same store — they never mix. Asking "how do competitors create urgency?" and "how do we?" are two different questions, so what each one learns stays its own, and a run reads back only what was formed under the question it is currently asking. Everything else in the store is kept exactly as it was and named in LEARNING_REPORT, so you can see it was left alone rather than merged. (Examples saved before this Actor recorded questions cannot be attributed to one, so they are preserved and reported, never read into a run.)

CODING_REPORT — in the run's key-value store

A receipt for the classification pass as a whole: pages_in_evidence, classifications, abstentions, and the dimensions read. When nothing matched it also carries a nothing_matched line saying so plainly, and what to do about it — "nothing matched" is a result, and you should be able to tell it apart from a run that quietly did nothing at all.

ambiguous_between — the second deliverable, and often the more interesting one. Some copy genuinely does two jobs in one sentence. When a passage sits between two of your labels, the run does not pick one and quietly move on; it records the pair and how many passages fell there:

"ambiguous_between": {
"authority / social_proof": 4,
"authority / urgency": 3,
"risk_reversal / social_proof": 3,
"scarcity / urgency": 2
},
"ambiguous_total": 12

Read that as a map of where the copy is doing double duty — "four passages are simultaneously appealing to expertise and to the crowd." A probabilistic classifier structurally cannot hand you this: it would have chosen one label and reported a confidence number, and the ambiguity — which is the actual finding about the writing — would have been the thing it destroyed. Here nothing is guessed and nothing is dropped. It is the same shape as the egress firewall's leak report, one Actor over: a receipt of what did not pass, and exactly why.

Field notes

  • question — the question you asked, in your own words, travelling with the row. So anyone who receives your data can tell what it was gathered to answer — evidence collected for one question is not evidence for a different one, and without this on the row there is no way for a downstream check to know. If you'd rather it didn't travel, don't allowlist it when you chain the firewall behind this Actor.
  • lens — which lens produced the row. A tag over your labels and trigger phrases, not their contents, so a proprietary lens stays proprietary while a client can still confirm that every row in an audit came from the frame you agreed on — and that a second audit used the same one. Sharpening a lens with Learn from this run does not change it; editing a label or a phrase does, because that is a different question being asked.
  • review_needed — reserved. A classification is emitted only after it clears internal cross-examination, so rows that ship read false; the field is there so a future review-flagged row has a home rather than changing shape on you.
  • images_alt_coverage — a 0–1 fraction, not a percentage.
  • provenance — see below. It is not an ID you need to use; it's a seal you can check.

How a label is decided

Worth understanding, because it explains both the accuracy and the silences.

Trigger phrases are a candidate-generator, not a verdict. A marker match pulls the surrounding sentence in as a candidate; the engine then decides the label by checking that candidate two independent ways at once. A label is emitted only when both readings agree and nothing else competes for it. Where they disagree, the engine abstains — no row, no charge for a finding that didn't hold.

Each passage is read against your lens — not against the rest of the batch. A passage is classified against the lens's own evidence: its trigger phrases and any examples you've accrued (see Learning from your own runs). What else happens to be in the same run doesn't change how a given passage is read — which is exactly what makes the result order-independent and reproducible. Two consequences worth expecting:

  • Depth comes from accruing, not from batch size. More pages in one run means more passages classified — each against the same lens — not a sharper read of any one of them. The lens gets sharper across runs when you switch on Learn from this run, so it stands on your own confirmed material instead of only the phrases you typed. A big one-off crawl doesn't buy accuracy; a lens you keep teaching does. (A single page classifies fine — there's no minimum.)
  • A page can match a trigger phrase and still yield no label. That is the check working, not a miss. "Learn more" contains earn; "only for existing customers" isn't scarcity. A quiet page is a real result — plain informational copy genuinely isn't running persuasion moves.

In seo and code modes the results land once the run has gathered and classified, rather than trickling in per page; scrape mode streams as it goes.

That's why precision-over-recall here is a capability rather than an apology: when it does speak, the label has already survived cross-examination.


Provenance

Every row carries a provenance tag computed over the row's own fields:

  • dove0:… — a SHA-256 content hash. Anyone can recompute it; it proves the row hasn't been edited.
  • dove1:… — an HMAC, when you supply a lens key. Only holders of the key can verify, which is what lets you prove to a client that a row came from your audit and nothing was touched after.

Every row also carries a second, independent tag — egress (the night seal, to provenance's day): computed separately over the same row, so a tampered row has to defeat both, and either can be checked on its own.


Modes

ModeWhat you get
SEO + marketing framing (seo)The full on-page signal set plus the persuasion-framing classification, each with its exact quote.
Multimodal scrape (scrape)Clean, readable page text. Turn on Multimodal to render the page and transcribe on-screen text — useful for JavaScript-heavy pages or text baked into images.
Classify against your own lens (code)Bring your own labels + trigger phrases and classify any list of pages against them. General-purpose, deterministic tagging with evidence quotes.

Input reference

{
"mode": "seo", // "seo" | "scrape" | "code"
"startUrls": [
{ "url": "https://example.com" },
{ "requestsFromUrl": "https://example.com/sitemap.xml" } // …or a .txt list of URLs
],
"multimodal": false, // render + transcribe on-screen text
"proxyConfiguration": { "useApifyProxy": true },// route fetches through Apify Proxy (default on)
"followLinks": false, // crawl: follow same-site links found on each page
"maxDepth": 1, // link-hops past your starting URLs
"lensName": "marketing_seo_v1", // built-in framing lens (seo/code modes)
"maxItems": 50, // hard stop, and your spend cap
"timeoutSecs": 30
}

urls accepts a plain list of strings as an alternative to startUrls. A requestsFromUrl entry is fetched and expanded before the run starts — XML sitemaps (including one level of sitemap index) and plain text lists both work.

Bringing your own lens (code mode)

This is the actual product, so it gets a proper explanation rather than a code sample.

What a lens is

A lens is the set of distinctions you want found in a body of text, written down so a machine can apply them the same way twice. It has exactly two levels:

  • a dimension — one question you are asking of every passage ("what pricing signal is this?")
  • its labels — the mutually exclusive answers to that question (free_tier, enterprise, …)

and each label carries trigger phrases — the words that make a sentence worth looking at for that label.

{
"mode": "code",
"startUrls": [{ "url": "https://example.com/pricing" }],
"lensJson": {
"matching_modes": {
"pricing_signals": { // ← the dimension: one question
"free_tier": { "lexical_markers": ["free plan", "free tier", "no credit card"] },
"enterprise": { "lexical_markers": ["contact sales", "custom pricing", "enterprise"] }
}
}
}
}

You can have several dimensions at once. Each is answered independently, so pricing signal and tone don't compete — a sentence can be free_tier on one and informal on the other.

Trigger phrases gather — they do not decide

This is the one thing worth internalising, because it is the opposite of how keyword tools work.

A marker match does not produce a result. It pulls the surrounding sentence in as a candidate, and then the engine reads that sentence against every label's evidence in two independent directions and only emits a label when both directions land on the same one — and when the words are actually there. A sentence your marker caught can absolutely come back with a different label, or with none.

So: write markers generously. A loose marker costs you nothing but a discarded candidate. A missing marker costs you a finding you'll never see, because a sentence nothing gathers is never read.

Prefer phrases to words

The one way to write a lens badly is to build it from short, common words.

save, results, more, faster, grow are ordinary English. A label made of them will gather API documentation, navigation menus and release notes, and — because those words genuinely are on the page — the engine will honestly report what it found. It is not wrong; your lens asked for it.

no credit card, contact sales, cancel anytime, limited time are phrases people only write when they mean the thing. Two or three words is usually the difference between a lens that reads a site and a lens that reads the whole internet.

Make the labels within a dimension genuinely different

Labels compete inside their dimension. If two of them overlap in meaning, passages will keep landing between them — and the run will tell you so, by name, in ambiguous_between. That is not a failure to fix by choosing a winner; it is the lens reporting that the distinction you drew isn't cleanly present in the copy. Either the two labels want merging, or the copy really is doing two jobs at once, and both of those are findings.

How to tell whether your lens is any good

Run it on twenty pages you already understand and open CODING_REPORT:

what you seewhat it meanswhat to do
lots of classifications, quotes look rightthe lens fits the materialkeep going
almost nothing, high abstentionsmarkers too narrow, or too few pagesadd phrasings; run more pages
findings on junk — menus, code, boilerplatemarkers are common wordsmake them phrases
high ambiguous_total, same pair repeatingtwo labels aren't distinct heremerge them, or accept both

Then read the evidence_quote on a handful of rows. Every one is a real sentence from a real page, so you are checking your own definitions against reality — not auditing a model's opinion.

Don't start from a blank page

You do not have to invent phrases out of your head, and you shouldn't — the copy you're studying already contains them.

Run scrape mode over ten or twenty of your target pages first and read the text it returns. The phrases people actually use will be sitting right there, in their words rather than your guess at their words, and a lens built from real copy beats one built from imagination on the first run. Then classify with code mode, read CODING_REPORT, and adjust. Two or three rounds of that is usually the whole job.

The lens sharpens itself, if you let it

A brand-new lens knows only the phrases you typed. Switch on Learn from this run and each run keeps the passages it confirmed, in a key-value store on your account, and reads them back next time — so the lens stops depending on your original guesses and starts standing on your own material. See Learning from your own runs above. Nothing is retained by us and you can delete it at any time.

Keeping a lens proprietary: ship it AES-256-GCM encrypted (encrypt_lens.py) and supply the passphrase as the secret input LENS_KEY. It decrypts in memory at runtime and never sits in the clear — your methodology stays your secret, even from us.


Output reference

seo mode:

{
"url": "https://example.com",
"title": "Example — Get Started Today",
"title_length": 30,
"meta_description": "Try it free, cancel anytime.",
"canonical": "https://example.com/",
"og": { "title": "Example", "description": "...", "image": "https://.../og.png", "type": "website" },
"twitter": { "card": "summary_large_image", "title": "Example" },
"h1_count": 1,
"word_count": 812,
"images_total": 14,
"images_with_alt": 11,
"images_alt_coverage": 0.786,
"links_internal": 42,
"links_external": 9,
"top_keywords": [{ "term": "pricing", "count": 12 }, { "term": "free", "count": 8 }],
"marketing_framing": [
{ "dimension": "persuasion_framing", "label": "risk_reversal",
"evidence_quote": "Try it free, cancel anytime.", "review_needed": false },
{ "dimension": "persuasion_framing", "label": "urgency",
"evidence_quote": "Get started today.", "review_needed": false }
],
"marketing_framing_counts": { "risk_reversal": 1, "urgency": 1 },
"provenance": "dove0:9f2c…"
}

code mode:

{ "dimension": "pricing_signals", "label": "free_tier",
"evidence_quote": "Start on the free plan, no credit card required.",
"source_url": "https://example.com/pricing", "source_title": "Pricing",
"review_needed": false, "provenance": "dove0:1a7b…" }

Pricing

About $0.01 per page. That is the whole answer for almost every run — a 100-page audit is roughly $0.70, a 1,000-page audit roughly $7. The detail below explains where that comes from and why you are never charged for a guess.

You pay only for what verifiably happened, and the bill is one number you can work out before you run: a small fixed run fee plus a flat per-page price. Two meters:

MeterWhen it firesPrice
run-startedOnce per run — boot, integrity self-check, browser spin-up (the fixed startup cost, priced as itself)$3.00 / 1,000
page-processedOnce per page fetched and fully analyzed — SEO signals and its marketing-framing classifications, each with a verbatim evidence quote$7.00 / 1,000

That's the whole meter — no per-result charge you can't predict. Your total is

run-started + (pages × page-processed)
, so a page costs the same whether it yields six framing labels or none. The $7.00 / 1,000 figure Apify shows in the header is your per-page price — the real driver of the bill — on top of a one-time $3.00 / 1,000 run fee that rounds to nothing on any real run. Because the engine is LLM-free, that per-page price buys a full scrape, SEO audit, and evidence-backed framing analysis for well under a cent.

What a real run costs

Concrete, and knowable before you press run:

  • Quick 3-page check — 1 run + 3 pages ≈ $0.02
  • 100-page SEO audit — 1 run + 100 pages ≈ $0.70
  • 1,000-page audit — 1 run + 1,000 pages ≈ $7.00

A page the engine abstains on — nothing it can quote — costs exactly the same as any other: you're paying for the read, not for findings, so a guess can never appear on your bill and ambiguous copy never inflates it.

How this price was derived

Priced by derivation, not by market — no reference to anyone else's prices. The formula: measured cost ÷ platform share + a stated stewardship wage. From a real measured run (2026-07-22): 2 pages + 4 classifications cost $0.0078 total — a marginal page cost of roughly $0.002–0.003. After the platform's ~20% share, cost-recovery is about $0.003 a page; the rest of the $0.007 is an openly-stated stewardship wage for building and keeping an engine that reads meaning without guessing. Classifications aren't metered separately — they cost ~nothing to emit, so the work, and the price, is the page. The launch is priced once as itself, so a 3-page check and a 3,000-page audit each pay the true shape of their own cost — no cross-subsidy. And one thing you are never charged for: a guess — the engine abstains on ambiguous copy rather than inventing a classification to bill you for. When the meters change, the price is re-derived and this section updated.


Deterministic vs. an AI analyzer — a structural difference, not a slogan

Most tools that classify web copy at scale run on an LLM. That buys fluency — and it gives up three guarantees you can't get back, not because any vendor is careless, but because a probabilistic model structurally can't make them:

  • Reproducibility. Ask a model the same thing twice and the answer can drift. Ask this engine, and the same pages yield the same result — no sampling, no temperature, nothing to drift. The order you list your URLs in makes no difference either: the same set of pages returns byte-identical output however it's shuffled, which is checked on every build. (Two runs can still differ when a page changes or a dynamic page renders differently — that's the input moving, never the engine.)
  • Evidence. A model can return a confident label with nothing on the page behind it, and can't prove it didn't. Here, every label carries the verbatim sentence that triggered it, or it isn't emitted — a guess has no path onto your bill.
  • Privacy. A model has to send your content off to judge it. This engine is LLM-free, so your pages never leave the container and never reach any AI provider. Privacy by construction, not policy.

And because there's no model call per result, runs stay cheap and fast across thousands of pages — which is what keeps the per-page price where it is.

None of that calls an AI tool dishonest. It's the difference between what a probabilistic method can offer and what a deterministic one can guarantee — and we'd rather you check it than take our word. Run the same audit twice and diff the output. Trace any label to its quoted sentence. Watch the network and see nothing leave. The difference isn't a claim you have to trust; it's one you can verify from where you sit.


Limits & FAQ

I got zero framing labels. Is it broken? Check CODING_REPORT in the key-value store — it tells you whether the run classified anything at all. Either the copy genuinely isn't running persuasion moves, or a candidate failed cross-examination, or the run was too small to read across (see How a label is decided). Silence is a result here, and it costs the same as a finding.

How many pages should I run for framing? Run as many as you want classified — each page is read against your lens independently, so there's no minimum and a single page classifies fine. What sharpens the lens is repetition over time with Learn from this run on, not the size of any one batch.

Can it audit my whole site? Two ways — see Covering a whole site. Point it at your sitemap (exact, no surprises), or switch on Follow links.

Will crawling blow up my bill? No — Max pages is a hard stop and the meter is one charge per page fetched, crawled or not. The queue itself is bounded by the remaining budget, so the run can never line up work it isn't allowed to do.

Why did a page match a trigger phrase but produce nothing? Markers gather candidates; they don't decide. See How a label is decided.

When should I turn on multimodal? For JavaScript-heavy pages, or text baked into images — where it reads what non-rendering scrapers can't see at all. It renders every page, so it's slower; leave it off for standard HTML.

Will every page fetch? Most will. Fetches route through Apify Proxy by default (datacenter — switch to Residential under Proxy for the hardest sites), which clears most bot-protection and geo-walls. What still can't be read: pages behind a login or paywall (there's no credential to give), and pages rendered entirely in JavaScript — turn on Multimodal for those. A page that can't be fetched comes back as a visible error row, never a silent gap. (Datacenter proxy is included on most Apify plans; residential adds Apify's own proxy usage to your account, billed by Apify at cost — separate from this Actor's per-page price.)

Does my content go to an AI provider? No. The engine is LLM-free and makes no model calls; pages are processed entirely in-container.

What about tone, imagery, and implication? Not classified. If there's no sentence to quote, it stays silent rather than guess.

Scraping etiquette: when crawling, the Actor honors robots.txt for every link it discovers itself. URLs you supply are fetched as asked — respecting each target site's terms of service is yours to weigh.


Feeding a grounding check

Every row this Actor emits already matches what the Grounded Answer Door reads: a span of source text under evidence_quote, its source_url, the question it was gathered under, and the lens that produced it. Point that Actor at this one's dataset and it will check a written answer sentence by sentence against exactly this evidence — no glue, no mapping.


Pairs with: Dataset Egress Firewall — and why the pair is the point

Analyzing pages for a client and shipping them the result? Chain Dataset Egress Firewall behind this Actor: name the fields allowed to leave, and it emits only those — with a leak report you can hand to the client as a receipt. Deterministic and AI-free on the same terms as this one.

The pairing is worth more than convenience, because the two govern different grains of the same rule.

The firewall decides which fields may leave. That is complete structurally and, alone, blind to content — a key pasted inside an allowed description has no field name to catch.

This Actor governs content: it emits only what a declared label licensed, with the sentence that earned it. Text nothing licensed is simply never in the output — not because something recognised it as dangerous, but because unlicensed content is never a candidate for emission at all.

raw records ──► this Actor: only licensed, quoted content ──► firewall: only named fields ──► out

A scanner enumerates what you fear, so it is always one novelty behind. An allowlist names what you mean, so what you did not name was never a candidate. Running that at the content grain and the field grain is what makes the guarantee structural instead of probabilistic. Neither Actor claims it alone.


Servicing of terms

We flipped the label on purpose: not terms that govern the service — a service that keeps its terms. Here they are, short enough to actually read: your question is yours (kept verbatim, never rewritten or interpreted, and it defines which of your accrued examples a run reads — a different question keeps its own ground, and nothing of yours is ever blended with anything else); your data stays yours (processed in-container, never retained after the run, never sold, never trained on, never shown to any AI provider); no rights are claimed over your inputs or outputs beyond mechanically running the job you asked for; you pay only for receipted events — never for a guess; you can leave anytime with nothing held. We ask one thing back: point it only at pages you're entitled to read, and respect each site's terms of service.

Support

Questions, or a custom lens for your use case? Open an issue on the Actor's Issues tab.