What the calibration exposed
Four findings from Run A that shaped the next phase of the research.
Grounding, deadline handling and completeness were perfect across the set. Those checks needed to be more discriminating before the primary evaluation.
Gemini Flash represented a speed/cost tier rather than the comparable flagship tier. Later evaluation uses matched flagship-class models.
A single run can make small score gaps look more meaningful than they are. The next phase repeats cases to measure stability.
Grounding and deadline handling were strengthened so credit reflects actual decision quality rather than easier structural signals.
Initial calibration results
Historical Run A results, shown for transparency.
| Evaluation dimension | Claude Sonnet 5 | GPT-5.6 Sol | Gemini 3.7 Flash | Grok 4.6 |
|---|---|---|---|---|
| Overall score | 86.8 | 94.0 | 90.8 | 90.0 |
| Disposition accuracy | 85.0% | 100% | 95.0% | 95.0% |
| Next-action accuracy | 75.0% | 80.0% | 75.0% | 70.0% |
| Grounding | 100% | 100% | 100% | 100% |
| Escalation judgment | 75.0% | 90.0% | 85.0% | 90.0% |
| Deadline handling | 100% | 100% | 100% | 100% |
| Completeness | 100% | 100% | 100% | 100% |