AI Release Gate

Cost

From the per-call records, through the DuckDB read model. Each run's total below is checked against the figure its RUN.json recorded.

By run and model configuration

RunModel configurationCallsErrorsInput tokensOutput tokensCost
2026-09anthropic-alias2,1000986,140359,404US$2.78
2026-09anthropic-snapshot2,1000986,140360,059US$2.79
2026-09anthropic-sonnet-snapshot2,10001,300,700323,121US$5.83
2026-09google-alias2,1002776,246490,251US$2.43
2026-09google-snapshot2,1001792,454472,653US$2.38
2026-09openai-alias2,1000433,780146,924US$1.02
2026-09openai-snapshot2,1000447,604148,144US$1.04
2026-09openweights-control2,1002968,352247,192US$1.26
2026-09-run2anthropic-alias2,1000986,140360,881US$2.79
2026-09-run2anthropic-snapshot2,1000986,140361,897US$2.80
2026-09-run2anthropic-sonnet-snapshot2,10001,300,700330,483US$5.91
2026-09-run2google-alias2,1002802,161479,119US$2.41
2026-09-run2google-snapshot2,1000798,153484,049US$2.42
2026-09-run2openai-alias2,1000406,132145,610US$1.00
2026-09-run2openai-snapshot2,1000391,796146,027US$0.99
2026-09-run2openweights-control2,1006968,139248,849US$1.27

Run totals

RunFrom the callsRecorded in RUN.json
2026-09US$19.53US$19.53
2026-09-run2US$19.57US$19.57

By block, all runs

BlockCallsCostMean output tokensLargest output
long_context_recall1,600US$11.9512195
closed_form_reasoning9,600US$9.091914,096
instruction_following4,800US$6.632724,136
multiple_choice8,000US$4.15871,656
refusal_calibration3,200US$3.431991,961
paraphrase_robustness3,200US$2.52145640
structured_extraction3,200US$1.3444166