RevCycleAIAI
Revenue Cycle Intelligence · Published Daily
RevCycleAI

Research

Independent benchmarks and original analysis for revenue cycle leaders.

We test AI systems, examine market structure, and publish transparent research on the workflows reshaping healthcare revenue cycle. DenialsBench is our first benchmark program; future RevCycleAI Research will extend beyond denials into other high-impact RCM workflows.

Featured research program: DenialsBench. A transparent benchmark series testing frontier AI performance on realistic revenue cycle denial scenarios. The current release is Run C, our repeatability gate ahead of the 100-case primary evaluation.
Featured program DenialsBench
Current study Run C — Repeatability
Framework DenialsBench v0.1
Next 100-case primary evaluation

DenialsBench · Run C repeatability results

Mean and variance across three independent 20-case runs.

Read the full analysis →
How to read this: a smaller standard deviation means the score moved less across repetitions. Mean scores should not be treated as a durable ranking when observed ranges overlap.
Repeatability measureGPT-5.6 SolClaude Opus 5Gemini 3.1 Pro PreviewGrok 4.6
Run scoresrep 1 · rep 2 · rep 388.0 · 92.5 · 90.089.0 · 93.3 · 90.391.3 · 90.8 · 88.087.0 · 87.0 · 85.3
Mean score0–10090.1790.8790.0386.43
Standard deviationlower = more stable1.841.801.450.80
Observed range88.0–92.589.0–93.388.0–91.385.3–87.0
API / parse failures0000
Critical failures0000
Grounding violations0000
Mean latency3.33s3.40s6.04s15.94s
Median latency3.15s2.95s5.59s13.78s
P95 latency4.88s4.94s9.73s28.17s
What Run C says so far: GPT-5.6 Sol, Claude Opus 5 and Gemini 3.1 Pro Preview are tightly clustered at this sample size. Their ranges overlap substantially, so the small mean differences should not be treated as a durable leaderboard. Grok scored lower on mean but was the most stable across the three repetitions.

Four useful takeaways

What the repeatability test changed about our interpretation.

0.70
GPT vs Claude mean gap

Too small to call a meaningful leader from Run C. Their observed score ranges overlap.

0.80
Grok standard deviation

Grok was the most repeatable model across the three repetitions.

90.03
Gemini clean rerun mean

Three clean repetitions with zero API/parse, critical or grounding failures.

~3.4×
Grok speedup at low reasoning

A separate low-reasoning check scored 86.5 with 4.75s mean latency versus 15.94s at default-high reasoning.

Research progression

DenialsBench is being published as a sequence, not a one-off leaderboard.

Run A — Calibration
Historical framework-validation run that exposed saturated dimensions, unmatched model tiers and the need for repeatability testing.

View Run A →

Next — 100-case primary evaluation
The first candidate public comparative scorecard will evaluate the full 100-case primary bank under the revised protocol.

Read methodology context →