Judge calibration
An LLM judge is used only where it has been shown to agree with a human. Each judge below was compared with 480 answers labelled by hand, blind to the model and to the judge. Below a kappa of 0.6 on a task the gate refuses to use that judge for it.
| Judge | Task | Status | Cohen's kappa | Krippendorff's alpha | Sensitivity | Specificity | Raw agreement | Above saying the majority |
|---|---|---|---|---|---|---|---|---|
google-judge-mid | faithful | refused | 10.8% (1.2 to 21.7) | 6.7% (-4.1 to 18.7) | 87.5% (84.3 to 90.4) | 55.6% (22.2 to 88.9) | 86.9% | -11.2% |
google-judge-mid | complete | licensed | 91.4% (87.5 to 95.0) | 91.4% (87.5 to 95.0) | 100.0% (99.2 to 100.0) | 89.0% (84.0 to 93.3) | 96.2% | +30.2% |
openai-judge-small | faithful | refused | 6.6% (-1.0 to 15.4) | 1.4% (-7.5 to 11.4) | 85.4% (82.0 to 88.3) | 44.4% (11.1 to 77.8) | 84.6% | -13.5% |
openai-judge-small | complete | licensed | 79.0% (72.9 to 84.8) | 78.9% (72.7 to 84.8) | 98.4% (96.8 to 99.7) | 76.7% (69.9 to 82.8) | 91.0% | +25.0% |
Raw agreement can look excellent on a lopsided set: when the human said yes 98% of the time, a judge that always says yes agrees 98%. The last column is raw agreement minus that, which is why a judge with 87% agreement can be refused.