Evidence
We publish what it actually does, not what we assume.
Every operator carries a catalogue assumption — the rate at which it is expected to bank its
outcome. Every deployment reports what it really did. Both numbers are in the console you are given, side by
side, with the count behind them. A vendor unwilling to show you the second number is quoting you the first.
| Operator | Owns | Assumed | Observed across clients |
Range | Outcomes behind it |
Illustrative of the reporting format. Your own figures replace these
from the first week of a pilot.
Four test suites
Golden cases written by people who used to do the work. Adversarial cases — confused, hostile and evasive
customers, and prompts designed to move the agent. Policy cases drawn from your never-list and your
regulator. Regression cases mined from production.
Policy is not a percentage
A single policy failure holds the release, whatever the rest of the suite says. No change reaches your
accounts without clearing the gate, and the gate is raised for anything carrying money or a regulator.
Clawbacks are labels
Because we are paid on evidenced outcomes, every banked result is a positive label and every clawback a
negative one. The test set that matters most is the one nobody has to write — and it grows every month
you run.
Shadow, then canary, then full. A new version runs alongside the live one
without sending anything, then takes a small share of real traffic, then everything — with the previous version
kept warm for rollback. On a squad that contacts customers we do not skip the first step, on any account,
for any deadline.