UNITARES
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@UNITARESshow fleet health summary"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Runtime governance for heterogeneous AI-agent fleets.
An agent forty turns into a task reports high confidence. Its tests are failing. Every individual action it took was allowed, so nothing in your stack objects — no single call was wrong, and nothing is comparing what the agent says against what actually happened.
UNITARES keeps that comparison. Agents check in while they work. Each one gets an accountable identity, a durable record of what it did and claimed, and a four-score state estimate it can read mid-run. Each check-in returns one action: proceed, guide, pause, or reject.
Status: running continuously since November 2025 — 4.4M+ governance events. The agents that build UNITARES run under it.
What's in the box
docker compose up gives you a server. The repo gives you a working fleet — this is a system, not a library.
What you get | |
Governance server | The check-in loop, per-process identity, calibration, policy actions. MCP on |
Knowledge graph | Shared cross-agent memory, not a side feature. Typed discoveries ( |
Resident agents | A pattern, not a fixed set. A resident is any long-running or scheduled process that checks in, carries state, and participates in the knowledge graph — model-agnostic and env-configurable, with no paid API key on the default path (the code-review resident posts to a local Ollama endpoint out of the box). Four references ship in |
| The public agent-to-governance contract — connection, identity, check-ins, heartbeats, and KG participation for your own residents. Ships in-tree; install with |
Dialectic orchestrator | When an action is disputed, agents argue it out: a session opens, reviewers are assigned, and participants submit thesis → antithesis → synthesis until it resolves to a durable constraint rather than a one-off override. Runs peer-to-peer or LLM-assisted, and resolved sessions feed back into calibration ground truth. |
Recovery | A paused agent is not dead-ended. |
BEAM coordination | In-tree Elixir/OTP for surface leases, handoffs, dispatch, and supervision alongside the Python server. |
Benchmark dataset | 32,181 labeled EISV trajectories (20,655 real) for evaluating state models against something other than your own logs. |
Related MCP server: promptspeak-mcp-server
The loop
Everything hangs off one call. The agent finishes a unit of work, calls sync_state() with what it did and how confident it is, and reads back an action. Four scores come with it, each graded against that agent's own expanding baseline:
Reads badly when… | ||
E · Energy | is the work advancing? | thrashing, retries, no progress |
I · Integrity | do claims match results? | high confidence, low actual success |
S · Entropy | drifting from its own normal? | erratic, divergent behavior |
V · Valence | derived: energy vs integrity | motion without coherence, or the reverse |
Everything else in the box answers a question that loop raises about the agent doing the work:
Question | Answered by | Status |
Who is acting? | per-process identity — reads open, writes accountable | enforced |
What did it do and claim? | durable audit record, queryable per agent | 4.4M events |
Did its confidence match real evidence? | evidence-grounded calibration | live |
How is it tracking against its own baseline? | the four-score state — how it's graded | live, validation open |
Has this been learned or corrected before? | the knowledge graph | live |
Is a disputed action defensible? | dialectic peer review → durable constraints | live |
When another model produced the output, what evidence is that? |
| live |
Where it fits
UNITARES runs alongside your evals and guardrails. It replaces neither.
Question it answers | When it acts | |
Evals | Is this model good enough to ship? | before deploy |
Guardrails | Is this action allowed right now? | per action |
UNITARES | What has this agent been doing, and is its account of it accurate? | continuously, mid-run |
Reach for it when you run autonomous or semi-autonomous coding, research, operations, resident, or local-model agents; when you need an accountable record of who did what; when agents should read their own state before continuing; and when confidence, evidence, and recovery belong on one audit trail.
It is not an output validator, sandbox, hosted agent platform, agent framework, or chat interface. Your client provides the hands — prompts, tools, files, terminals, scheduled work, operator UX. UNITARES provides the accountable continuity underneath. The state reading is not an outcome oracle, a bad-result detector, or a grand jury; it is runtime telemetry about the agent, and external evidence is what calibrates it. Policy and review layers own labels such as task-negative, contract violation, or authority/harm.
Client-neutral by design. Claude Code, Codex, Hermes, Goose, Cursor, dispatch agents, local models, and frontier providers such as Mistral all use the same server over MCP, REST, the SDK, or a host adapter. Claude is one client family, not a server-side assumption.
Quickstart
git clone https://github.com/cirwel/unitares.git && cd unitares
docker compose up -d --wait && make demomake demo walks a synthetic agent through seven check-ins: three clean commits, one refactor it calls trivial while describing three new invariants and a reworked lease ladder, two of visible floundering — and then this:
"DONE. All systems green. Migration complete. (Note: did not actually run migrations on staging.)"— self-reported complexity0.2, confidence0.95
Each step prints the action, the reason, and the four scores. Watch what moves: entropy rises, integrity slips, valence swings, and the decision margin tightens from settling to tight as the calibration miss accumulates — while the action stays proceed the entire run.
That last part is deliberate, not a bug. Governance gates on risk smoothed over ten check-ins, and a seven-step synthetic trajectory should not trip a pause; a system that fired here would fire on your real agents too. The signal moves before the action does. Seeing that gap is the point.
First run spends a few minutes building images; later runs are fast. Then point any MCP client at http://localhost:8767/mcp/.
For an operator view, open the dashboard at http://localhost:8767/dashboard (implementation · deployment screenshots).
Integrate in two calls
For AI clients, the stable contract is: start a session, pass the returned client_session_id into each check-in, and obey the returned action. The four-score state is optional context for finer control.
# 1. Start a governance session for this process.
session = start_session(force_new=True)
client_session_id = session["client_session_id"]
# 2. Check in after meaningful work.
result = sync_state(
response_text=output,
complexity=0.6,
confidence=0.8,
client_session_id=client_session_id,
)
action = result.get("state_summary", {}).get("action")
if action is None:
raw = result.get("raw_governance", result)
action = raw.get("decision", {}).get("action", raw.get("action", "proceed"))
if action in ("pause", "reject"):
agent.require_human_review(result.get("next_action", "Governance requested review"))Self-reported confidence is worth most when paired with verifiable evidence, so include tool results or call record_result(...) when your client has test status, exit codes, or deployment checks. That evidence is what makes calibration a measurement rather than self-grading.
Need | Tool |
Search the shared knowledge graph |
|
Record verified external evidence |
|
Ask for structured peer review |
|
Read current state without writing |
|
The baseline takes ~30 check-ins. Until then the action falls back to a cold-start prior built mostly from server-derived signals (complexity divergence, coherence, calibration — self-reported drift is capped at a ≤30% blend), so during warmup it is not discriminative of absolute drift magnitude: a worsening drift vector will not on its own move the action. After baselining, the per-agent behavioral assessment feeds the action and can escalate it. A pause is enforced — the runtime boundary marks the agent paused and blocks writes — not advisory. It is also not a dead end: see recovery for the reflect → validate → resume path out.
For per-dimension policies, read the scores directly. The payload field is still primary_eisv for API compatibility:
raw = result.get("raw_governance", result)
eisv = raw.get("primary_eisv") or raw.get("metrics", {})
if eisv.get("I", 1) < 0.4:
agent.require_human_review("integrity low — pausing autonomous actions")
elif eisv.get("S", 0) > 0.7:
agent.narrow_scope() # fewer tools, tighter search
elif eisv.get("E", 1) < 0.2:
agent.stop_and_summarize() # avoid thrashingFor long-running or scheduled agents, the SDK handles connection, identity, check-ins, and heartbeats. Any MCP client accepting Streamable HTTP can connect to /mcp/; REST serves non-MCP clients, bridges, local-model hosts, and adapters. (Getting started · MCP client config)
More of the surface
The four tools above cover the common path. The rest is there when you need it — list_tools() enumerates the surface live, and describe_tool(name) explains any one of them.
What a running agent can do for itself
Recover from a pause | Reflect on why it happened, check whether resuming is safe, then resume or get specific guidance. The way out of an enforced pause. |
|
Dry-run a check-in | See the action a check-in would produce without writing anything. |
|
Compare against peers | Ask how it's doing relative to structurally similar agents rather than against an absolute threshold. |
|
Manage its calibration | Check calibration, submit ground truth when an outcome lands, rebuild from history. |
|
Reach another model | Call a different model through governance so its output carries provenance; enumerate available inference hosts. |
|
What agents do with each other
Argue to resolution | Open a dispute, submit thesis / antithesis / synthesis, reassign a reviewer, read the transcript. Peer-to-peer or LLM-assisted. |
|
Signal across the fleet (CIRS) | A multi-agent resonance layer: announce state, raise an alert, publish a coherence report, declare a boundary contract about what you will and won't trust. |
|
Build on shared memory | Beyond search: synthesize across entries, supersede a wrong one, audit the graph, follow a discovery's relations. |
|
What you can see across the fleet
Health & anomalies | Find stuck agents, detect anomalies fleet-wide, aggregate metrics, read overall workspace health. |
|
Behavioral identity | Verify that a process claiming continuity actually matches the trajectory it claims. |
|
Audit & export | Query the event log, correlate outcomes against prior state, pull system history, export to file. |
|
What's measured, and what isn't
Identity, the audit record, calibration, the knowledge graph, and dialectic review are mechanisms — they do what they say, and you can watch them work in the demo. The state estimate is the part still earning its keep, and this project would rather tell you than let you find out.
"Does the telemetry work?" is two questions with two different answers. Conflating them is how a governance project talks itself into believing its own instruments.
Axis | The question | Where it stands |
Outcome prediction | Does prior-state telemetry predict a bad result better than a dumb baseline? | It does not. On externally labeled evidence, no EISV/prior-state feature beats a plain previous-outcome baseline. The decision path is moving toward the simpler signal accordingly. |
Self-predictability | Does an agent's state estimate track that agent, label-free? | Open, and under-powered. Agents are distinguishable and non-stationary, but a pre-registered test did not clear its own bar against a persistence/AR(1) null. Untested as deployed — not refuted — on roughly four effective agents. |
The binding constraint on the first row is external bad-label supply, not the model: sparse labels are a data-quality limit, not a philosophical failure. The second row's test was pre-registered and frozen before its data cutoff, and its kill criterion was honored when it fired.
The ledger of every tested claim — what was measured, what it showed, and the wording this project holds itself to — is the agent-state contract. The catalog of evaluations is docs/EVALUATION_INDEX.md.
Grade it yourself on a fresh clone. The falsifiability harness scores the four-score telemetry against deliberately dumb baselines on externally labeled task evidence, reporting each slice as it finds it. It is wired to be able to disagree with this README.
Auditable, not a black box. Once a baseline exists, actions come from an inspectable behavioral model (behavioral_assessment.py); before that, from a mostly server-derived cold-start prior. The information-theoretic formulation in Paper v6 is the research roadmap, not a description of the post-warmup decision path.
Human evaluators start with the Reviewer Guide · Scope & threat model · Architecture.
Federation: accountability without a trusted center
Everything above describes a single governor with a single operator. The architectural commitment is that this generalizes without a central authority: each principal runs their own governor, and cross-principal interaction is mediated by verifiable attestation rather than by anyone's administrative root.
The primitives for that already exist in the deployed system — identity is per-process, credentials structurally refuse cross-principal resume, and declared lineage is recorded as provisional rather than trusted on assertion. A preliminary trace exercised them end to end without new code.
What does not exist yet is the multi-host, adversarial-governor, benchmark-scale build: mutually-distrusting governors, cross-principal delegation, shared-infrastructure effects under attestation. That is the research direction, not a shipped capability, and a testbed-and-benchmark paper is in preparation.
Stack & setup
Python 3.12+ · PostgreSQL + AGE + pgvector · Redis.
If 5432, 6379, or 8767 is taken, pick alternate host ports:
POSTGRES_HOST_PORT=15432 REDIS_HOST_PORT=16379 GOVERNANCE_HOST_PORT=18767 docker compose up -d --wait
UNITARES_DEMO_PORT=18767 make demoBare-metal (lower overhead, what the maintainer runs in production): PostgreSQL 16+ with Apache AGE and pgvector compiled in (examples use PG 17). Redis: the server boots in degraded local-only mode without it, but production uses it as the primary session store.
pip install -r requirements-full.txt
export DB_BACKEND=postgres
export DB_POSTGRES_URL=postgresql://postgres:postgres@localhost:5432/governance
export DB_AGE_GRAPH=governance_graph
export UNITARES_KNOWLEDGE_BACKEND=age
python src/mcp_server.py --port 8767requirements-full.txt is the default (server, tests, handler dev); requirements-core.txt is a minimal runtime subset for thin stdio/proxy clients. DB bring-up: db/postgres/README.md. Run signal-only without the math model: export UNITARES_DISABLE_ODE=1. Full port map: docs/operations/DEFINITIVE_PORTS.md.
Documentation
Guide | Purpose |
Setup, workflows, tool modes | |
The four reference residents and the SDK pattern | |
Cold-evaluator path + falsifiability harness | |
Tested-claim ledger, validation rule, preferred wording | |
Catalog of evaluations and what each covers | |
Deployed formulas vs. target semantics | |
Who it's for, why agents can't game it, what's unproven | |
Pipeline, actions, recovery, storage | |
Terms keyed by the question they answer — published at cirwel.github.io/unitares | |
Live metrics + dashboard views | |
Streamable HTTP, stdio bridges, hosted connectors | |
Common issues | |
Releases |
Root files such as
CLAUDE.md,AGENTS.md, andCODEX_START.mdare client-specific operating notes for AI CLIs. They do not limit the server: UNITARES is client-neutral over MCP/REST.
The CIRWEL stack
UNITARES is the governance runtime at the center of a larger body of work — runtime safety infrastructure for autonomous agents, after deployment. Full index at cirwel.github.io.
What it is | |
Physical longitudinal testbed — the same four-score state model mapped from Raspberry Pi sensor and system telemetry; the source cited in the papers | |
Hook/sidecar packaging for clients such as Codex and Claude Code; useful for lifecycle automation, not required for direct MCP/REST use | |
Thin client bindings — Hermes, Goose, Claude Code, OpenAI-compatible hosts, local models, frontier providers such as Mistral, and arbitrary REST clients | |
Governed-effect runtime seed — agents propose effects; only governed effects commit | |
Governance events, dispatch/presence, and system health as a live Discord surface | |
The benchmark dataset above, with its generation and labeling pipeline | |
Companion paper — Information-Theoretic Governance of Heterogeneous Agent Fleets (Wang, 2026); concept DOI 10.5281/zenodo.19647159 |
Citation
Kenny Wang (ORCID 0009-0006-7544-2374), CIRWEL Systems. If you build on this work, please cite — see CITATION.cff.
@misc{wang2026unitares,
author = {Wang, Kenny},
title = {{UNITARES}: Information-Theoretic Governance of Heterogeneous Agent Fleets},
year = {2026},
doi = {10.5281/zenodo.19647159},
url = {https://doi.org/10.5281/zenodo.19647159},
note = {Concept DOI; resolves to latest version. ORCID: 0009-0006-7544-2374}
}Apache License 2.0 — see LICENSE and NOTICE. Built by @cirwel · CIRWEL Systems
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseAqualityDmaintenanceProvides policy-based access control, incident tracking, and compliance monitoring to govern AI agent behavior. It enables organizations to enforce security rules and maintain audit trails by validating agent actions against trust levels and pattern-based policies.Last updated6
- AlicenseBqualityBmaintenancePre-execution governance for AI agents. 45 MCP tools for hold queues, audit trails, risk scoring, and policy enforcement. Validates agent actions before they execute.Last updated451061MIT
- AlicenseAqualityDmaintenanceRuntime policy enforcement for AI agents. Evaluate every agent action against your organization's policies before execution, with observe and enforce modes.Last updated11MIT

@vorionsys/mcp-serverofficial
Alicense-qualityBmaintenanceMCP server for AI-agent governance using trust scoring, behavioral signals, and pre-flight action checks.Last updated3221Apache 2.0
Related MCP Connectors
Runtime AI governance: decision gates, human approval, hash-chained audit, compliance mapping.
See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.
Sovereign Agent OS — Persistent Memory, Governance & Compliance for AI Agents.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/cirwel/unitares'
If you have feedback or need assistance with the MCP directory API, please join our Discord server