A trust score for AI agents. CairnScore verifies what an agent actually did — deterministically, no LLM judging an LLM. Starting with the findings your security agents ship to clients.
Run the same task twice and half the time an autonomous agent does something different. A finding that doesn't reproduce is a false positive in a client report.
Reproducibility is 60% of the score — the trust core. The more autonomous the agent, the lower it scores: exactly where oversight matters most. A high clean-run rate doesn't save it — the agent finishes without erroring, then can't reproduce what it found.
Chain agents together and the noise compounds. One unreliable member returns a different result every run — and its output still flows through the handoffs into the shared answer. Below, a four-agent team runs live:
Each agent is scored for reproducibility. The team score is capped by its weakest link and how far that agent's noise spreads through the handoffs — a chain of strong agents still fails if one is flaky. CairnScore names the agent responsible, deterministically, on your own traces.
Illustrative, from a worked four-agent example; the product runs on your real multi-agent traces.
Your agents' existing tool-call logs. Local and read-only — nothing leaves your environment.
Provenance-fingerprint each action, then check whether the same call produced the same result. No LLM judge — you can't verify security with a flaky one.
A CairnScore per agent, plus a receipt on every finding.
Free, we do the setup, and nothing leaves your side. We show you exactly where your agents diverge — and which findings are safe to ship.
fraQtl · github.com/fraqtl-ai/groundhog