AI Release Gate

Judge calibration

An LLM judge is used only where it has been shown to agree with a human. Each judge below was compared with 480 answers labelled by hand, blind to the model and to the judge. Below a kappa of 0.6 on a task the gate refuses to use that judge for it.

JudgeTaskStatusCohen's kappaKrippendorff's alphaSensitivitySpecificityRaw agreementAbove saying the majority
google-judge-midfaithfulrefused10.8% (1.2 to 21.7)6.7% (-4.1 to 18.7)87.5% (84.3 to 90.4)55.6% (22.2 to 88.9)86.9%-11.2%
google-judge-midcompletelicensed91.4% (87.5 to 95.0)91.4% (87.5 to 95.0)100.0% (99.2 to 100.0)89.0% (84.0 to 93.3)96.2%+30.2%
openai-judge-smallfaithfulrefused6.6% (-1.0 to 15.4)1.4% (-7.5 to 11.4)85.4% (82.0 to 88.3)44.4% (11.1 to 77.8)84.6%-13.5%
openai-judge-smallcompletelicensed79.0% (72.9 to 84.8)78.9% (72.7 to 84.8)98.4% (96.8 to 99.7)76.7% (69.9 to 82.8)91.0%+25.0%

Raw agreement can look excellent on a lopsided set: when the human said yes 98% of the time, a judge that always says yes agrees 98%. The last column is raw agreement minus that, which is why a judge with 87% agreement can be refused.