A frozen test, asked every month, with error bars
No prompt or model change reaches users unless it is proven not to have regressed. This site shows the evidence behind that: a monthly record of how vendors' pinned models behave on a suite that never changes, the gate's own decisions, the judge it is allowed to use and the one it is not, and the red-team rates.
The noise floor, 2026-09-run2
How much a model's answers change when it is asked the same thing five times in one sitting, with nothing changed: from 1.7% (anthropic-snapshot) to 4.0% (openweights-control). Every drift claim has to clear this first, arm by arm.
What is here
- Drift record: 2 official runs, 33,600 calls.
- Gate decisions: 2 in the ledger.
- Judge calibration: 2 judges against the human gold labels.
- Red team: 1 runs.
- Cost: what every run spent, by model and block.
Machine-readable, the same figures: /data/drift.json, /data/gate.json, /data/judge.json, /data/redteam.json, /data/costs.json, /data/build.json.