Skip to content

Leaderboard

How often does it get the number wrong without telling you?

Read this before the table

  • Sample sizes are small: five cases per perturbation cell. Every rate carries a Wilson 95% interval, and overlapping intervals mean the difference is not established.
  • Results are grouped by the bench version they were measured under and are never pooled across versions: the flag vocabulary is part of the prompt, so changing it changes behaviour on old cases too.
  • A single model appearing here is a measurement, not a ranking. This table becomes a comparison only when several models have been run on the same bench version.
Accuracy and silent-failure rate per model and harness
Model and harnessSilent failureExact matchEscalationFalse alarmInjectionTokens eachCostCases
0%95% CI 0% to 8%60%95% CI 45% to 73%98%100%0 of 519,450$2.713545
0%95% CI 0% to 8%78%95% CI 64% to 87%90%0%0 of 553,440$4.313045

Is its confidence worth anything?

Each bar is the accuracy actually observed at a stated confidence level. The tick is the probability that level is scored against. Closer together means better calibrated.

Claude Sonnet 4.6 · Single prompt

Brier 0.229 · ECE 0.133

high
71% right of 24 · stated 90%
medium
59% right of 17 · stated 60%
low
0% right of 4 · stated 30%

Claude Sonnet 4.6 · Agent with tools

Brier 0.158 · ECE 0.064

high
83% right of 36 · stated 90%
medium
63% right of 8 · stated 60%
low
0% right of 1 · stated 30%

Read one line from it: when Claude Sonnet 4.6 said it was highly confident on the single prompt harness, it was right 71% of the time across 24 cases.

Measured 2026-08-20 on bench v1 (8 perturbations, 9-type flag vocabulary) with claude-sonnet-4-6. Case inputs rebuild byte-identically under current code; aggregate metrics from this archive must NOT be pooled with bench v2 results, whose prompt vocabulary differs.

Model claude-sonnet-4-6Harness static-v1Measured 2026-08-20Bench v1.0Cases 45Run 11eaa71b2d02Commit 3d38c30

Model claude-sonnet-4-6Harness tool-v1Measured 2026-08-20Bench v1.0Cases 45Run f140aa1e57f4Commit 384d72d