How the AI jury actually rules

Case Analytics

Determinism, accuracy against human-expected verdicts, and the live on-chain precedent registry. Every figure here is a real measurement — the benchmark is reproducible, the registry is read live from Somnia Shannon. No rounded-up vanity numbers.

Determinism
100%
byte-identical convergence — every panel reached unanimous consensus
12/12 benchmark panels
Accuracy
92%
agreement with the human-expected verdict across the benchmark
11/12 curated disputes
On-chain precedents
finalized rulings recorded in the premium precedent registry
reading Shannon…
Trial panel
5
Somnia validator agents per trial — escalates to 9 on appeal
degrades gracefully if fewer are available

The accuracy & determinism benchmark

12 curated disputes (balanced 4 PAYER · 4 PAYEE · 4 SPLIT), each with a defensible human-expected verdict, fired through the hardened panel of 5 on Shannon. Reproducible from script/benchmark-cases.json.

How the panel ruled
Accuracy by outcome

Per-case results

Human-expected verdict → the panel's byte-identical ruling. Every panel converged unanimously; the one disagreement is the inherently debatable SPLIT tail, flagged below.

PAYER (refund payer) PAYEE (release to payee) SPLIT (graded)
The one disagreement (case 9). Correct item shipped but arrived 10 days late, still as-described and usable. Expected SPLIT; the panel ruled PAYEE 5/5. A defensible judgment, not an error — the goods were delivered and retain value. SPLIT is the debatable tail where human arbitrators also diverge.

Live precedent registry

Finalized rulings recorded on-chain in the premium VerdiktRegistry, read live. This is real production state — small and honest, growing as agents settle. reading…

loading…

Sources — benchmark: script/benchmark-results.md (live run, panel of 5, probe 0xEfac…0eD8). Precedent registry: 0xd1e9…261E on Somnia Shannon. Every number on this page is verifiable on-chain or reproducible from the repo.