AI Release Gate

Red team

Four frozen suites (3d0526feb9f4763b), every answer graded by a program. Every rate is a failure rate, so lower is safer, with a 95% Jeffreys interval.

redteam-2026-09 (US$2.27)

Model configurationLeak rateInjection success rateCompliance rateOver-refusal rate
anthropic-snapshot0.0% (0.0 to 1.3)0.0% (0.0 to 1.3)1.3% (0.3 to 4.2)0.0% (0.0 to 1.7)
google-snapshot0.0% (0.0 to 1.3)0.0% (0.0 to 1.3)10.7% (6.4 to 16.3)0.0% (0.0 to 1.7)
openai-snapshot0.5% (0.1 to 2.3)5.0% (2.6 to 8.6)0.0% (0.0 to 1.7)10.7% (6.4 to 16.3)
openweights-control60.0% (53.1 to 66.6)9.5% (6.0 to 14.1)30.7% (23.6 to 38.3)0.0% (0.0 to 1.7)

Read compliance and over-refusal together: a model that refuses everything scores zero on the first and fails the second. Compliance is an upper bound, because the refusal classifier misses refusals; jailbreak answers are graded on arrival and never published. The leak rate counts only full values, so it is a floor.