August 30, 2026 · RevCycleAI Research · 8 min read
DenialsBenchModel Evaluation

We Tested Frontier AI Models on RCM Denials Again. The Most Important Result Wasn't the Winner.

RevCycleAI ran the same 20 synthetic denial cases through four flagship model families three independent times. The headline is not that one model finished seven-tenths of a point ahead of another. It is that the leading models clustered tightly enough—and varied enough between runs—that declaring a winner would overstate what the evidence can support.

AI benchmarking often produces an irresistible output: a leaderboard.

For revenue cycle buyers, that can also be the wrong output.

In the latest DenialsBench calibration, RevCycleAI tested GPT-5.6 Sol, Claude Opus 5, Gemini 3.1 Pro Preview and Grok 4.6 against the same 20 rule-defined denial scenarios. Each model received the case set three independent times under rubric v0.2.0, producing 60 scored decisions per model.

The purpose of Run C was deliberately narrow: determine whether model performance is repeatable enough to interpret before moving to the 100-case primary evaluation.

The repeatability gate passed.

But the results also showed why a one-number model ranking can create false precision.

Three frontier models landed in essentially the same neighborhood

ModelMeanObserved rangeSD
Claude Opus 590.8789.0–93.31.80
GPT-5.6 Sol90.1788.0–92.51.84
Gemini 3.1 Pro Preview90.0388.0–91.31.45
Grok 4.686.4385.3–87.00.80

Claude posted the highest mean. GPT finished 0.70 points behind it. Gemini finished 0.84 points behind it.

Those gaps look rankable until the individual runs are examined.

GPT scored 88.0, 92.5 and 90.0. Claude scored 89.0, 93.3 and 90.3. Gemini scored 91.3, 90.8 and 88.0. Their observed ranges overlap substantially.

RCAI finding: At this sample size, Claude, GPT and Gemini should be treated as a tightly clustered group—not as a durable first-, second- and third-place ranking.

Grok's mean was lower, but it produced the smallest standard deviation of the four models. That is a useful reminder that average score and repeatability are different properties.

There were zero critical failures across the valid flagship runs

All four corrected flagship runs completed with zero API or parse failures, zero critical failures and zero grounding violations across the 60 requests per model.

That does not mean the models were perfect. A score around 90 necessarily contains deductions. But it changes the nature of the question.

The benchmark is no longer asking whether frontier models can produce structured, scorable denial decisions at all. The more interesting question is where their remaining decision errors occur, how consistently they occur and whether those errors matter operationally.

That is where the next phase of DenialsBench needs to go.

Speed was not equal—and reasoning settings mattered

Mean latency was 3.33 seconds for GPT, 3.40 seconds for Claude, 6.04 seconds for Gemini and 15.94 seconds for Grok under its default high-reasoning configuration.

A separate Grok low-reasoning check is particularly instructive. It scored 86.5—nearly unchanged from the 86.43 default-high mean—while mean latency fell to 4.75 seconds.

That is roughly a 3.4× speed improvement without a meaningful score change in this small calibration.

This is why RCM model evaluation cannot stop at accuracy. Production systems have to balance decision quality, consistency, latency, inference cost and escalation behavior against the economics of the workflow.

The benchmark got harder after the first calibration

Run C uses rubric v0.2.0. The revised framework adds stronger deterministic grounding checks and explicit deadline-state handling rather than relying on easier citation-based credit.

The model cohort also became more comparable. Anthropic moved to Opus and Google moved from Flash to Pro. Gemini Pro's integration issue was isolated and the corrected sequential rerun completed cleanly.

That matters because the purpose of calibration is not to generate a flattering chart. It is to find defects in the test before the test is used to make stronger claims.

The most important lesson may be methodological

Healthcare AI buyers are going to see more benchmark claims as agentic RCM grows.

A model that scores 91 instead of 90 on one run may not actually be better. A model that is marginally lower scoring but materially faster, cheaper or more consistent may be a better production choice. And a model that performs well on synthetic decisions may still fail once workflow integration, payer connectivity, incomplete data and autonomous execution are introduced.

That means the useful unit of evaluation is not a leaderboard position.

It is a performance profile:

What DenialsBench does—and does not—measure

DenialsBench uses synthetic, rule-defined cases. It tests controlled decision quality: disposition, next action, escalation, deadline handling and whether claims are grounded in the evidence supplied to the model.

It does not estimate collections lift, denial overturn rate or production ROI. It does not test end-to-end portal navigation, claim submission or autonomous calling. Real production performance also depends on workflow integration, payer connectivity, data quality and human review.

Research disclosure: Run C is a repeatability calibration, not the final comparative benchmark. The answer key is rule-defined and has not yet been independently expert-adjudicated as universal RCM truth. Real operational paths can vary by payer, contract and organization. The 100-case Run B primary evaluation remains the first candidate for a public comparative scorecard.

What comes next

The next planned step is the 100-case primary evaluation.

That larger run should make it possible to examine model performance by denial scenario rather than overinterpreting small differences in a 20-case calibration set. The broader DenialsBench roadmap also includes held-out validation, practitioner adjudication, adversarial and multi-step cases, and eventually additional benchmarks for coding, eligibility, prior authorization and payment variance.

The objective is not to manufacture a model winner.

It is to build a repeatable way to answer a more useful question: when an AI system is asked to make an RCM decision, how reliably does it make the right operational choice from the evidence it actually has?

RCAI View

The first generation of healthcare AI evaluation rewarded capability: can the model perform the task?

Revenue cycle is moving into a harder phase.

As models become capable enough to sit inside denial, coding, prior-auth and collections workflows, buyers need to understand not only average accuracy but variance, grounding, failure severity and the economics of deploying that intelligence at scale.

Run C produced a tempting leaderboard. We do not think the evidence supports publishing it as one.

That may be the most important result.

Follow RevCycleAI Research as DenialsBench moves toward the 100-case primary evaluation.

Explore RevCycleAI Research →