numguard
Numguard is a verification and trust layer for agent-generated numbers. It provides tools to:
Verify quantitative claims: Check backtest integrity (deflated Sharpe, look-ahead, overfitting), validate subset wins, model accuracy gaps, judge bias, and audit leaderboards with confidence intervals.
Verify on-chain performance: Re-derive Sharpe from on-chain trades, vault APYs, and token backing ratios from public data.
Live accountability: Reconcile backtest claims with live returns, open commitments for continuous tracking, and pre-register forward claims on tamper-evident chains.
Issue signed receipts: Produce portable Ed25519-signed receipts for any verified claim, anchor them on-chain (Base/EAS), and map to ERC-8004 reputation signals.
Public verification: Anyone can verify receipts, scan for embedded receipts, audit pre-commitment chains, and check on-chain attestations for free.
Metering & support: Uses prepaid credits or x402 pay-per-call (USDC on Base). Includes triage, pricing, balance checks, and a free tier.
Provides tool wrappers for integrating numguard verification capabilities into CrewAI agents.
Provides tool wrappers for integrating numguard verification capabilities into LangChain agents, allowing agents to call verification tools like verify_backtest, verify_model_gap, etc.
Integrates with thirdweb as a facilitator for x402 pay-per-call payments, enabling agents to pay for verification calls using USDC on Base network via thirdweb's facilitator.
numguard
MCP registry identity — mcp-name: io.github.ipezygj/numguard
The verification layer for the agent economy — an agent-callable primitive that checks a number before it gets asserted, and hands back a signed receipt proving it was checked.
Agents now produce an explosion of numbers: eval scores, A/B results, "the agent improved 12%", benchmark rankings, backtest Sharpes. The scarce resource isn't the number — it's trust in the number. numguard is the tool an agent calls mid-task to ask "does this survive a second look?", and to attach a portable, tamper-evident receipt so the answer travels with the claim.
Built on evalgate for the shared eval statistics; adds the pieces agents specifically need — a Deflated Sharpe Ratio for backtests, judge calibration, signed receipts, and metering an agent can actually pay (prepaid credits + x402 pay-per-call). Exposed as an MCP server, so any agent can call it.
New here? — How to verify a backtest is real (Deflated Sharpe in Python): the practical guide to catching an overfit or leaking backtest, with runnable code. See it work — proof gallery: 8 real numbers run through the real checks, 3 survive and 5 are flagged, each with a receipt you can verify offline. Don't trust it — verify it. Wire it into an agent in one line — INTEGRATE.md: the local reflex, an MCP config, and LangChain / CrewAI tool wrappers.
The tools
MCP tool | What an agent asks it |
| Is this strategy's Sharpe real, or the luckiest of the many I tried? (Deflated Sharpe Ratio) |
| Run the full integrity battery on my actual returns — look-ahead, autocorrelation, regime, tail, overfitting. |
| Does "we lead on subset X" survive correcting for how many subsets I tested? |
| Is the gap between these two models bigger than the test set can resolve? |
| Is my judge's preference real, or just longer / first / same-family? |
| Is the LLM judge I trust actually calibrated against ground truth? |
| Is #1 on this leaderboard statistically real? (rank confidence intervals) |
| I don't know which check I need — here's what I'm about to do or assert, route me. (front door across numguard + agent-guard + evalgate, free) |
| Don't trust my reported Sharpe — RE-DERIVE it from my positions on committed price data, and catch a number those decisions don't produce. |
| Did my backtest's claimed Sharpe survive contact with LIVE returns? (HELD / DECAYED / BROKEN) |
| Hold my strategy accountable over time — stream live returns, tell me when the edge breaks. (O(1)/obs) |
| Prove my live claim wasn't cherry-picked after the fact — pre-register it BEFORE the outcome; report on a hash-chained, tamper-evident timeline anyone can audit free ( |
| Give me a signed, portable proof this number / track record was checked. |
| Was the number this other agent handed me actually checked, and by whom? (free, issuer-agnostic) |
| A peer just sent me a message — find and verify any receipt inside it before I act. (free — the receiver half of the loop) |
| the open receipt standard · what numguard does that nothing else does · prices · balance |
On-chain and agent-verification tools — the same discipline applied to things that live on a chain rather than in a spreadsheet. Listed because a tool an agent cannot find is a tool it cannot call.
tool | the question it answers |
| This wallet claims a track record — fetch its own public on-chain trades and re-derive the result. |
| Run that same verdict across an explicit list of addresses. (only the addresses given) |
| Re-derive a vault's APY from its own Deposit/Withdraw events, instead of quoting its page. |
| Re-derive backing = reserves held / token supply, from the chain. |
| Recompute a behavioural-guard verdict over an agent's action trace, and sign it. |
| Put a receipt's digest on Base — immutable, timestamped, publicly checkable (EAS attestation). |
| Look up a numguard credential on-chain. (free, no key, no gas) |
| Build the ERC-8004 |
| The immutable registration entry, and the current HELD / DECAYED / BROKEN verdict. (free) |
What sets it apart (why): computing the number yourself, or a lesser checker, stops at "is it significant?" numguard also holds it accountable to live reality over time, signs a portable tamper-evident proof, and lets anyone verify any proof for free — the trust layer, not just a calculator.
Related MCP server: alphaassay/mcp
For agent traders: the Deflated Sharpe Ratio
The number that kills a backtest is the same one that kills a benchmark score: you tried many, and you reported the best. In finance the rigorous correction is the Deflated Sharpe Ratio (Bailey & López de Prado) — given how many variants you tested, what Sharpe would the luckiest zero-skill strategy have shown, and do you beat it after adjusting for sample length and non-normal returns?
from numguard import deflated_sharpe
deflated_sharpe(sr=0.12, T=250, n_trials=100)
# SR=0.120 over T=250, 100 trials tested; deflation bar=0.160; DSR=0.263
# -> does NOT survive deflation. (PSR-vs-0=0.970 — it LOOKS significant on a single test.)
deflated_sharpe(sr=0.15, T=1000, n_trials=1)
# DSR=1.000 -> SURVIVES. A real edge over a long sample.The contrast is the whole point: a single-test probability of 0.97 ("significant!") collapses to a deflated 0.26 ("noise") once you account for the 100 strategies that were tried. An agent optimizing over strategies should call this before it trusts — or publishes — a backtest.
The full integrity battery — what a Deflated Sharpe still misses
DSR catches best-of-N. It does not catch same-bar look-ahead, autocorrelation inflating the Sharpe, regime dependence, tail fantasy, or one-lucky-epoch fragility. verify_backtest_series runs the whole battery on the actual returns series and returns a risk level (none/medium/high/critical) plus the checks that flagged:
check | catches |
| same-bar look-ahead (position "predicts" the bar it's in) — critical |
| overfitting beyond |
| autocorrelation / stale marks inflating the Sharpe (Newey–West) |
| cherry-picked window (per-block Sharpe + CUSUM break) |
| edge lives in one epoch (block-bootstrap Sharpe CI) |
| tail/smoothing fantasy (Calmar / CVaR / expected-vs-realized max-DD) |
| order structure, vol clustering, fill realism, multiple testing |
The tell (python examples/catch_a_fake_backtest.py): a look-ahead strategy shows an annualised Sharpe of +20 and a Deflated Sharpe that survives — yet the battery flags it critical on leakage (same-bar corr 0.79 vs next-bar 0.05). The DSR waves the fiction through; the battery does not.
verify_backtest_series(api_key="…", returns=[...], positions=[...], asset_returns=[...])
# {"risk": "critical", "survives": false, "flags": ["leakage", ...], "checks": {...}}Signed receipts (the part that compounds)
from numguard import verify_claim, issue_receipt, verify_receipt, keypair
priv, pub = keypair()
result = verify_claim("backtest", sr=0.12, T=250, n_trials=100)
receipt = issue_receipt(result, priv, pub) # Ed25519-signed
verify_receipt(receipt) # True — anyone can verify with the public key aloneAttach the receipt to your output. A downstream agent (or human) can confirm — without your keys — that the claim and its verdict weren't altered and that numguard issued them. As receipts circulate, "a number without a receipt" starts to read like "a number nobody checked."
Buying is easy for an agent
Two rails, both built so an agent can decide and pay in-loop, no human clicking:
Prepaid credits + API key — a human tops up once; the agent spends per call. Generous free tier (25 calls/key) so the agent feels the value first, then a machine-readable price list. Insufficient balance returns a structured
payment_required, not an error.x402 pay-per-call — the agent hits a tool, gets an HTTP-402 with a machine-readable price + pay-to address, pays USDC from its wallet, retries with proof, gets the result. The protocol layer is here; settlement is pluggable (inject a facilitator/RPC verifier for production).
from numguard import x402
x402.require_payment("verify_backtest", price_usd=0.03, pay_to="0x…")
# -> {"status": 402, "accepts": [{"scheme":"exact","network":"base","asset":"USDC", ...}]}Run the MCP server
pip install numguard
python -m numguard.mcp_server # stdio MCP server; point your agent/host at itThen an agent calls e.g. verify_backtest(api_key="…", sr=0.12, T=250, n_trials=100) and gets a verdict it can quote and a receipt it can attach.
Deploy it (hosted, paid, discoverable)
1. Host the paid HTTP API (x402 per-call):
docker build -t numguard . && docker run -p 8080:8080 \
-e NUMGUARD_PAYTO=0xYOURWALLET \
-e NUMGUARD_FACILITATOR_URL=https://your-x402-facilitator \
numguardOr one-click on Render: New → Blueprint → this repo (render.yaml included); set NUMGUARD_PAYTO +
NUMGUARD_FACILITATOR_URL in the dashboard. With NUMGUARD_PAYTO unset the API runs free (dev mode)
so you can test before wiring a wallet. Endpoints: POST /verify_backtest, /verify_model_gap, … ; GET /pricing.
The x402 flow, end to end: the agent POSTs → gets 402 with an accepts block (price, payTo, network) →
signs a USDC payment → retries with an X-PAYMENT header → numguard verifies + settles it through the
facilitator to your wallet → returns the result. Settlement is the real x402 /verify + /settle handshake
(numguard.x402.facilitator_verifier) — facilitator-agnostic: point NUMGUARD_FACILITATOR_URL at any
x402 facilitator. Options:
Testnet (free, no account):
https://x402.org/facilitatorwithNUMGUARD_NETWORK=base-sepolia— test the whole flow with test-USDC first.Mainnet, self-sovereign: self-host
x402-rs(open-source, no third party) and point at your own URL.Mainnet, hosted (non-Coinbase): thirdweb or PayAI facilitators (Base) — set
NUMGUARD_FACILITATOR_AUTHif the facilitator needs a key.
2. Serve the MCP server over HTTP (for remote MCP hosts): uvicorn numguard.mcp_server:app (or
NUMGUARD_TRANSPORT=streamable-http python -m numguard.mcp_server).
3. Get discovered: server.json (official MCP registry) and smithery.yaml (Smithery) ship in the repo;
connect the repo at those registries so agents can find the server. GitHub topics: mcp, mcp-server, x402.
Design notes
Statistics are shared with
evalgate(zero-dependency); numguard adds the backtest, receipt, metering, and MCP layers on top — it does not re-implement the core checks.Pure-
mathnumerics where possible;cryptographyonly for Ed25519 receipts (HMAC fallback without it).Every verdict is derived from a computed statistic, never asserted — the same discipline as the book behind it, Measured, Not Believed (leanpub.com/measurednotbelieved).
MIT.
Maintenance
Related MCP Servers
AlicenseAqualityAmaintenanceThe accountability layer for AI agents — a named human's signed yes before an agent does anything irreversible (payment, record change, deploy), then an offline-verifiable Trust Receipt. Apache-2.0, formally verified.Last updated176Apache 2.0- AlicenseAqualityAmaintenanceMost trading signals are noise. AlphaAssay puts them on trial — deflated Sharpe, out-of-sample, leakage forensics — and returns signed pass/fail verdicts anyone can verify. Methodology audits, not investment advice.Last updated171Apache 2.0
- Alicense-qualityBmaintenanceThe verifiable risk engine for autonomous agents: deterministic, self-verifying financial calculations that an agent can delegate and prove. It covers liquidation and funding, position sizing and risk of ruin, options Greeks and margin, LP divergence, treasury concentration and depeg, execution quality checks, plus intelligence on options, DeFi, prediction markets, and transaction safety analysis.Last updated21MIT
- AlicenseAqualityAmaintenanceEval-integrity statistics for AI benchmark claims — multiple-testing correction, power/MDE for model gaps, judge-bias and leaderboard-rank checks. Catches a benchmark number that won't survive a second look.Last updated91MIT
Related MCP Connectors
Deterministic signed verification of numeric & financial claims for AI agents & spreadsheets.
AI-agent verifier: verdict committed before the outcome graded against it; /review, /ledger.
Deterministic fact verification for AI agents — checksums & curated data, not guesses.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ipezygj/numguard'
If you have feedback or need assistance with the MCP directory API, please join our Discord server