DenialsBench · Run C repeatability results
Mean and variance across three independent 20-case runs.
| Repeatability measure | GPT-5.6 Sol | Claude Opus 5 | Gemini 3.1 Pro Preview | Grok 4.6 |
|---|---|---|---|---|
| Run scoresrep 1 · rep 2 · rep 3 | 88.0 · 92.5 · 90.0 | 89.0 · 93.3 · 90.3 | 91.3 · 90.8 · 88.0 | 87.0 · 87.0 · 85.3 |
| Mean score0–100 | 90.17 | 90.87 | 90.03 | 86.43 |
| Standard deviationlower = more stable | 1.84 | 1.80 | 1.45 | 0.80 |
| Observed range | 88.0–92.5 | 89.0–93.3 | 88.0–91.3 | 85.3–87.0 |
| API / parse failures | 0 | 0 | 0 | 0 |
| Critical failures | 0 | 0 | 0 | 0 |
| Grounding violations | 0 | 0 | 0 | 0 |
| Mean latency | 3.33s | 3.40s | 6.04s | 15.94s |
| Median latency | 3.15s | 2.95s | 5.59s | 13.78s |
| P95 latency | 4.88s | 4.94s | 9.73s | 28.17s |
Four useful takeaways
What the repeatability test changed about our interpretation.
Too small to call a meaningful leader from Run C. Their observed score ranges overlap.
Grok was the most repeatable model across the three repetitions.
Three clean repetitions with zero API/parse, critical or grounding failures.
A separate low-reasoning check scored 86.5 with 4.75s mean latency versus 15.94s at default-high reasoning.
Research progression
DenialsBench is being published as a sequence, not a one-off leaderboard.
Run A — Calibration
Historical framework-validation run that exposed saturated dimensions, unmatched model tiers and the need for repeatability testing.
Next — 100-case primary evaluation
The first candidate public comparative scorecard will evaluate the full 100-case primary bank under the revised protocol.