Deterministic agent verification

Keep your AI agents honest.

A trust score for AI agents. CairnScore verifies what an agent actually did — deterministically, no LLM judging an LLM. Starting with the findings your security agents ship to clients.

Measured on public pentest-agent traces · your numbers run on your own data
The problem, measured

Autonomous agents reproduced only 51% of their actions — vs 80% with a human in the loop.

Run the same task twice and half the time an autonomous agent does something different. A finding that doesn't reproduce is a false positive in a client report.

Human in the loop
80%
consistent, run to run
Fully autonomous
51%
half its actions diverge
Score card — real agents, scored
Human-assisted agenttrust core: passing
Reproducibility
79.6%
CairnScore 64 / 100clean-run 64%
Fully autonomous agenttrust core: at risk
Reproducibility
50.8%
CairnScore 58 / 100clean-run 86%

Reproducibility is 60% of the score — the trust core. The more autonomous the agent, the lower it scores: exactly where oversight matters most. A high clean-run rate doesn't save it — the agent finishes without erroring, then can't reproduce what it found.

The multi-agent problem

A team is only as honest as its weakest agent.

Chain agents together and the noise compounds. One unreliable member returns a different result every run — and its output still flows through the handoffs into the shared answer. Below, a four-agent team runs live:

atom = an agent electrons = its outputs arc = a handoff core = the shared answer
reliable — tight orbit weak link — scatters the answer forming
Team CairnScore
0/100
Weakest link
Researcher · 36% reliable
How the team is scored

Each agent is scored for reproducibility. The team score is capped by its weakest link and how far that agent's noise spreads through the handoffs — a chain of strong agents still fails if one is flaky. CairnScore names the agent responsible, deterministically, on your own traces.

Illustrative, from a worked four-agent example; the product runs on your real multi-agent traces.

How it works
01
Point it at traces

Your agents' existing tool-call logs. Local and read-only — nothing leaves your environment.

02
Verify deterministically

Provenance-fingerprint each action, then check whether the same call produced the same result. No LLM judge — you can't verify security with a flaky one.

03
Get a score + receipts

A CairnScore per agent, plus a receipt on every finding.

✓ REPRODUCED✗ FLAKY — do not ship
Why security agents first
Existential
Findings go to clients. A pentest agent that flags a vuln on run 1 and misses it on run 3 puts a false positive in a client report — a liability, not an inconvenience.
Deterministic
You can't judge security with an LLM. The observability crowd catches failures with an LLM grading an LLM — itself non-deterministic. For findings that go to clients, that's disqualifying. CairnScore is deterministic.
Scales
It keeps up with the swarm. As agents run concurrently by the thousand, humans can't review at agent speed. Machine-speed verification is the only thing that does.

We run CairnScore on a slice of your traces.

Free, we do the setup, and nothing leaves your side. We show you exactly where your agents diverge — and which findings are safe to ship.

Email us to start View the tool

fraQtl · github.com/fraqtl-ai/groundhog