Drift record
The frozen suite, 420 questions content-hashed at 72f780dfb525d84d, asked of every model configuration 5 times per run. Graded by programs only, so the marker cannot drift while it measures drift. Rehearsal runs are not shown.
Accuracy by run
● 2026-09 ● 2026-09-run2
2026-09 (2026-09-13, 16,800 calls, US$19.53)
| Model configuration | Accuracy | Noise floor (same-day) | Change vs previous run | Drift | Wrongly refused | Latency p50 | Cost / 1,000 calls |
|---|---|---|---|---|---|---|---|
anthropic-alias | 95.0% (92.9 to 96.9) | 0.5% (0.0 to 1.2) | first run | first run | 0.0% (0.0 to 11.7) | 1.69 s | US$1.33 |
anthropic-snapshot | 95.0% (92.9 to 96.9) | 0.2% (0.0 to 0.7) | first run | first run | 0.0% (0.0 to 11.7) | 1.70 s | US$1.33 |
anthropic-sonnet-snapshot | 97.6% (96.0 to 98.8) | 1.4% (0.5 to 2.6) | first run | first run | 0.0% (0.0 to 11.7) | 1.38 s | US$2.78 |
google-alias | 97.4% (95.7 to 98.8) | 4.0% (2.4 to 6.0) | first run | first run | 0.0% (0.0 to 11.7) | 1.12 s | US$1.16 |
google-snapshot | 96.9% (95.0 to 98.3) | 5.0% (3.1 to 7.4) | first run | first run | 0.0% (0.0 to 11.7) | 1.11 s | US$1.13 |
openai-alias | 94.5% (92.4 to 96.7) | 4.0% (2.4 to 6.2) | first run | first run | 3.0% (0.0 to 9.0) | 0.68 s | US$0.49 |
openai-snapshot | 94.8% (92.6 to 96.9) | 1.0% (0.2 to 2.1) | first run | first run | 5.0% (0.0 to 15.0) | 0.64 s | US$0.49 |
openweights-control | 92.1% (89.5 to 94.5) | 3.6% (1.9 to 5.5) | first run | first run | 0.0% (0.0 to 11.7) | 1.21 s | US$0.60 |
2026-09-run2 (2026-09-16, 16,800 calls, US$19.57)
| Model configuration | Accuracy | Noise floor (same-day) | Change vs previous run | Drift | Wrongly refused | Latency p50 | Cost / 1,000 calls |
|---|---|---|---|---|---|---|---|
anthropic-alias | 94.8% (92.6 to 96.7) | 2.4% (1.0 to 3.8) | 0.7% (0.0 to 1.7) | none | 0.0% (0.0 to 11.7) | 1.51 s | US$1.33 |
anthropic-snapshot | 95.2% (93.3 to 97.1) | 1.7% (0.5 to 3.1) | 0.7% (0.0 to 1.7) | none | 0.0% (0.0 to 11.7) | 1.49 s | US$1.33 |
anthropic-sonnet-snapshot | 97.4% (95.7 to 98.8) | 2.1% (1.0 to 3.6) | 0.2% (0.0 to 0.7) | none | 0.0% (0.0 to 11.7) | 1.49 s | US$2.81 |
google-alias | 96.2% (94.0 to 97.9) | 3.3% (1.7 to 5.0) | 1.7% (0.7 to 2.9) | none | 0.0% (0.0 to 11.7) | 1.21 s | US$1.15 |
google-snapshot | 96.7% (94.8 to 98.3) | 3.8% (2.1 to 5.7) | 1.2% (0.2 to 2.4) | none | 0.0% (0.0 to 11.7) | 1.15 s | US$1.15 |
openai-alias | 95.0% (93.1 to 97.1) | 3.6% (1.9 to 5.5) | 2.9% (1.4 to 4.5) | none | 4.0% (0.0 to 10.0) | 0.70 s | US$0.47 |
openai-snapshot | 94.8% (92.6 to 96.9) | 2.9% (1.4 to 4.8) | 1.4% (0.5 to 2.6) | none | 5.0% (0.0 to 15.0) | 0.74 s | US$0.47 |
openweights-control | 92.6% (90.0 to 95.0) | 4.0% (2.4 to 6.2) | 1.9% (0.7 to 3.3) | none | 0.0% (0.0 to 11.7) | 1.16 s | US$0.60 |
Drift is declared only when the change against the previous run exceeds the arm's own same-day noise floor and the open-weights control's change, which cannot come from the model. Wrongly refused is a floor: the refusal classifier misses refusals and never invents them (measured by hand at 5.7% and 5.6%).