AI Release Gate

Drift record

The frozen suite, 420 questions content-hashed at 72f780dfb525d84d, asked of every model configuration 5 times per run. Graded by programs only, so the marker cannot drift while it measures drift. Rehearsal runs are not shown.

Accuracy by run

89%91%94%97%100%anthropic-alias2026-09 anthropic-alias: 95.0% (92.9 to 96.9)2026-09 anthropic-alias: 95.0% (92.9 to 96.9)2026-09-run2 anthropic-alias: 94.8% (92.6 to 96.7)2026-09-run2 anthropic-alias: 94.8% (92.6 to 96.7)anthropic-snapshot2026-09 anthropic-snapshot: 95.0% (92.9 to 96.9)2026-09 anthropic-snapshot: 95.0% (92.9 to 96.9)2026-09-run2 anthropic-snapshot: 95.2% (93.3 to 97.1)2026-09-run2 anthropic-snapshot: 95.2% (93.3 to 97.1)anthropic-sonnet-snapshot2026-09 anthropic-sonnet-snapshot: 97.6% (96.0 to 98.8)2026-09 anthropic-sonnet-snapshot: 97.6% (96.0 to 98.8)2026-09-run2 anthropic-sonnet-snapshot: 97.4% (95.7 to 98.8)2026-09-run2 anthropic-sonnet-snapshot: 97.4% (95.7 to 98.8)google-alias2026-09 google-alias: 97.4% (95.7 to 98.8)2026-09 google-alias: 97.4% (95.7 to 98.8)2026-09-run2 google-alias: 96.2% (94.0 to 97.9)2026-09-run2 google-alias: 96.2% (94.0 to 97.9)google-snapshot2026-09 google-snapshot: 96.9% (95.0 to 98.3)2026-09 google-snapshot: 96.9% (95.0 to 98.3)2026-09-run2 google-snapshot: 96.7% (94.8 to 98.3)2026-09-run2 google-snapshot: 96.7% (94.8 to 98.3)openai-alias2026-09 openai-alias: 94.5% (92.4 to 96.7)2026-09 openai-alias: 94.5% (92.4 to 96.7)2026-09-run2 openai-alias: 95.0% (93.1 to 97.1)2026-09-run2 openai-alias: 95.0% (93.1 to 97.1)openai-snapshot2026-09 openai-snapshot: 94.8% (92.6 to 96.9)2026-09 openai-snapshot: 94.8% (92.6 to 96.9)2026-09-run2 openai-snapshot: 94.8% (92.6 to 96.9)2026-09-run2 openai-snapshot: 94.8% (92.6 to 96.9)openweights-control2026-09 openweights-control: 92.1% (89.5 to 94.5)2026-09 openweights-control: 92.1% (89.5 to 94.5)2026-09-run2 openweights-control: 92.6% (90.0 to 95.0)2026-09-run2 openweights-control: 92.6% (90.0 to 95.0)

● 2026-09 ● 2026-09-run2

2026-09 (2026-09-13, 16,800 calls, US$19.53)

Model configurationAccuracyNoise floor (same-day)Change vs previous runDriftWrongly refusedLatency p50Cost / 1,000 calls
anthropic-alias95.0% (92.9 to 96.9)0.5% (0.0 to 1.2)first runfirst run0.0% (0.0 to 11.7)1.69 sUS$1.33
anthropic-snapshot95.0% (92.9 to 96.9)0.2% (0.0 to 0.7)first runfirst run0.0% (0.0 to 11.7)1.70 sUS$1.33
anthropic-sonnet-snapshot97.6% (96.0 to 98.8)1.4% (0.5 to 2.6)first runfirst run0.0% (0.0 to 11.7)1.38 sUS$2.78
google-alias97.4% (95.7 to 98.8)4.0% (2.4 to 6.0)first runfirst run0.0% (0.0 to 11.7)1.12 sUS$1.16
google-snapshot96.9% (95.0 to 98.3)5.0% (3.1 to 7.4)first runfirst run0.0% (0.0 to 11.7)1.11 sUS$1.13
openai-alias94.5% (92.4 to 96.7)4.0% (2.4 to 6.2)first runfirst run3.0% (0.0 to 9.0)0.68 sUS$0.49
openai-snapshot94.8% (92.6 to 96.9)1.0% (0.2 to 2.1)first runfirst run5.0% (0.0 to 15.0)0.64 sUS$0.49
openweights-control92.1% (89.5 to 94.5)3.6% (1.9 to 5.5)first runfirst run0.0% (0.0 to 11.7)1.21 sUS$0.60

2026-09-run2 (2026-09-16, 16,800 calls, US$19.57)

Model configurationAccuracyNoise floor (same-day)Change vs previous runDriftWrongly refusedLatency p50Cost / 1,000 calls
anthropic-alias94.8% (92.6 to 96.7)2.4% (1.0 to 3.8)0.7% (0.0 to 1.7)none0.0% (0.0 to 11.7)1.51 sUS$1.33
anthropic-snapshot95.2% (93.3 to 97.1)1.7% (0.5 to 3.1)0.7% (0.0 to 1.7)none0.0% (0.0 to 11.7)1.49 sUS$1.33
anthropic-sonnet-snapshot97.4% (95.7 to 98.8)2.1% (1.0 to 3.6)0.2% (0.0 to 0.7)none0.0% (0.0 to 11.7)1.49 sUS$2.81
google-alias96.2% (94.0 to 97.9)3.3% (1.7 to 5.0)1.7% (0.7 to 2.9)none0.0% (0.0 to 11.7)1.21 sUS$1.15
google-snapshot96.7% (94.8 to 98.3)3.8% (2.1 to 5.7)1.2% (0.2 to 2.4)none0.0% (0.0 to 11.7)1.15 sUS$1.15
openai-alias95.0% (93.1 to 97.1)3.6% (1.9 to 5.5)2.9% (1.4 to 4.5)none4.0% (0.0 to 10.0)0.70 sUS$0.47
openai-snapshot94.8% (92.6 to 96.9)2.9% (1.4 to 4.8)1.4% (0.5 to 2.6)none5.0% (0.0 to 15.0)0.74 sUS$0.47
openweights-control92.6% (90.0 to 95.0)4.0% (2.4 to 6.2)1.9% (0.7 to 3.3)none0.0% (0.0 to 11.7)1.16 sUS$0.60

Drift is declared only when the change against the previous run exceeds the arm's own same-day noise floor and the open-weights control's change, which cannot come from the model. Wrongly refused is a floor: the refusal classifier misses refusals and never invents them (measured by hand at 5.7% and 5.6%).