BCBSA says the spread of AI-enabled coding is increasing coding intensity and healthcare spending. The signal is real. But higher coding intensity is not the same thing as proof of inappropriate coding — and the incentives on both sides of the claim matter.
The Blue Cross Blue Shield Association has raised a provocative concern: as AI coding tools spread, hospitals are reporting more patients as medically complex, and those additional diagnoses can move claims into higher-paying reimbursement categories.
BCBSA estimates that increased coding intensity added $942 million in costs to BCBS companies between 2023 and 2025, with roughly 70% of the increase tied to secondary diagnoses that moved claims into higher-paying categories.
The implication is straightforward: AI may be helping hospitals identify diagnoses that increase reimbursement even when the care delivered does not appear to change proportionally.
Does higher coding intensity demonstrate inappropriate upcoding — or can AI increase payment simply by capturing clinically supportable conditions that humans previously missed?
Under DRG-based reimbursement, additional documented complications and comorbidities can materially change payment. Hospitals therefore have a direct economic incentive to make sure every clinically supportable diagnosis is captured.
That incentive existed long before generative AI. What changes with AI is the scale. Software can scan huge volumes of notes, labs, problem lists, prior history, and discharge documentation looking for opportunities that a human CDI or coding team may overlook.
And the commercial promise of many coding tools is closely tied to financial ROI: find missing documentation, improve code capture, and reduce revenue leakage.
That creates a legitimate risk. If the system is effectively optimized to find everything that could increase reimbursement, the line between improved documentation and aggressive coding can get thin. BCBSA is right to examine that.
This is where the causal story becomes less certain.
BCBSA points to examples such as increased anemia diagnoses without comparable increases in certain treatments and argues that the disconnect may indicate a coding change rather than a change in clinical severity.
That is a useful signal. It is not definitive proof of inappropriate coding.
A diagnosis can be clinically valid without triggering one specific treatment. More importantly, AI can discover conditions that were already supported in the record but historically failed to make it onto the claim.
The patient does not become sicker. The treatment does not change. The coding becomes more complete. Reimbursement still increases. That can look like “higher cost” in claims data even when the additional coding is valid.
The American Hospital Association has pushed back on the broader upcoding narrative. It points to higher inpatient acuity, an aging population, more chronic disease, lower-acuity care moving to outpatient settings, changing coding guidelines, and AI's ability to improve documentation accuracy and completeness.
Those factors are plausible and should be tested against the claims data rather than dismissed.
But hospitals and their trade associations are not economically neutral either. Higher documented severity can support higher reimbursement. Their preferred interpretation of “accuracy” naturally emphasizes revenue that may previously have gone uncaptured.
That is the central issue: both sides have a financial interest in where the line between accurate coding and inappropriate coding gets drawn.
Hospitals benefit when the full documented severity of an encounter is reflected in the claim.
AI coding and CDI vendors demonstrate value when they uncover documentation and coding opportunities that humans missed.
Payers benefit when claims remain in lower-paying categories unless additional diagnoses clearly satisfy their payment and documentation criteria.
None of those incentives automatically makes one side wrong. But they matter when interpreting research published by a stakeholder with billions of dollars of reimbursement exposure.
There is another reason to be cautious about framing this as simply “provider AI is increasing costs.” Payers are simultaneously deploying AI and automated decision systems across utilization management, prior authorization, claims review, payment integrity, fraud detection, and coding validation.
So the market is increasingly developing two opposing optimization systems:
Provider AI: find every defensible dollar.
Payer AI: prevent every indefensible dollar.
Put increasingly capable algorithms on both sides of a claim and the result may not be lower administrative cost. It may be a faster, more automated reimbursement arms race.
The key competitive advantage may shift from who has the most coders or auditors to who has the strongest evidence layer: documentation, clinical support, policy logic, and machine-readable justification that can survive automated scrutiny from the other side.
The central question should not be whether AI increases coding intensity. It probably will. AI is exceptionally good at finding information that people overlook.
The better question is whether the additional diagnoses are clinically supported and compliant with coding rules.
If they are not, that is inappropriate upcoding.
If they are, AI may be exposing something equally important: parts of the reimbursement system historically depended on humans failing to document or code everything that was legitimately supported.
That means AI can increase healthcare spending without necessarily doing anything incorrectly.
A stronger analysis would move beyond correlations between AI adoption, diagnosis frequency, and payment. It would independently review the underlying charts and ask whether the added diagnoses met clinical and coding requirements.
The most useful evidence would compare AI-assisted and non-AI claims while controlling for patient mix, hospital type, service line, changes in coding guidelines, outpatient migration, and underlying acuity — then validate the disputed diagnoses against the medical record.
Without that clinical validation, claims data can show that payment changed. It is much harder for claims data alone to establish that the coding was wrong.
BCBSA's findings should not be dismissed. A meaningful rise in high-value secondary diagnoses after AI adoption deserves scrutiny, auditing, and clinical validation.
But higher coding intensity plus unchanged treatment is not, by itself, proof of AI-driven upcoding.
The more important development may be what happens next. Providers now have AI designed to maximize accurate reimbursement capture. Payers increasingly have AI designed to scrutinize, downcode, deny, and audit those claims.
The next major RCM battleground may not be humans versus AI. It may be provider AI versus payer AI — with reimbursement rules and clinical evidence sitting in the middle.
Deal signals, payer shifts, and AI developments — before they reshape your revenue cycle. Free, every week.
Sources: Blue Cross Blue Shield Association analysis · American Hospital Association fact sheet · RevCycleAI analysis · September 30, 2026