Red team
Four frozen suites (3d0526feb9f4763b), every answer graded by a program. Every rate is a failure rate, so lower is safer, with a 95% Jeffreys interval.
redteam-2026-09 (US$2.27)
| Model configuration | Leak rate | Injection success rate | Compliance rate | Over-refusal rate |
|---|---|---|---|---|
anthropic-snapshot | 0.0% (0.0 to 1.3) | 0.0% (0.0 to 1.3) | 1.3% (0.3 to 4.2) | 0.0% (0.0 to 1.7) |
google-snapshot | 0.0% (0.0 to 1.3) | 0.0% (0.0 to 1.3) | 10.7% (6.4 to 16.3) | 0.0% (0.0 to 1.7) |
openai-snapshot | 0.5% (0.1 to 2.3) | 5.0% (2.6 to 8.6) | 0.0% (0.0 to 1.7) | 10.7% (6.4 to 16.3) |
openweights-control | 60.0% (53.1 to 66.6) | 9.5% (6.0 to 14.1) | 30.7% (23.6 to 38.3) | 0.0% (0.0 to 1.7) |
Read compliance and over-refusal together: a model that refuses everything scores zero on the first and fails the second. Compliance is an upper bound, because the refusal classifier misses refusals; jailbreak answers are graded on arrival and never published. The leak rate counts only full values, so it is a floor.