‹ Back to Blog

Healthcare

Denial and Underpayment Detection in 2026: The Healthcare Data Layer Behind AI That Actually Recovers Revenue

Zach Shapiro

·

·

15 min read

Clinicians and administrators meeting to review revenue cycle performance in a hospital

TL;DR: Denials are the visible half of revenue leakage. Underpayments are the half that funds your competitors. A denied claim creates a work queue item with an owner and an appeal. An underpaid claim posts, closes, and looks like revenue, which is why underpayments run an estimated 1% to 11% of net patient revenue. Detecting one requires knowing what was done, what was billed, what the contract says should be paid, what was actually paid, and why the payer says the difference exists, across five systems that do not agree on what an encounter is. This guide covers 2026 denial and underpayment rates, why RCM AI pilots stall after go-live, what a healthcare data layer is, and a realistic 90-day sequence.

Denials are the visible half of the problem. Underpayments are the half that funds your competitors. A denied claim generates a work queue item, an owner, and an appeal. An underpaid claim posts, closes, and looks like revenue. Nobody works it, because nothing told anyone it was short. Analysis compiled by Becker's Hospital Review puts annual losses to underpayments at 1% to 11% of net patient revenue. At a health system with $400 million in net patient revenue, the low end of that range is $4 million a year that no work queue will ever surface.

The reason underpayment detection stays unsolved is not that the analytics are hard. It is that detecting an underpayment requires simultaneously knowing what was done, what was billed, what the contract said should be paid, what was actually paid, and why the payer says the difference exists. Those five facts live in five systems that do not agree on what a patient encounter is. That is a data layer problem, and it is why the AI pilot your team ran last year produced a good demo and no recovered dollars.

This guide covers the current numbers, why existing RCM tooling misses underpayments, why healthcare AI pilots stall after go-live, what a healthcare data layer is, and a realistic 90-day sequence. It is written for CFOs, revenue cycle leaders, practice administrators, and physician group owners.

Key takeaways

  • Initial denial rates have climbed to roughly 11.65%, near 12% industry-wide in 2026, meaning revenue is at risk on more than one in nine claims. Average denied amounts rose 14% in hospital outpatient and 12% in inpatient settings year over year.

  • Medicare Advantage is the sharpest case. A Health Affairs study covering about 30% of the MA market found initial denial rates of 17%, with 57% of those denials ultimately overturned on appeal. Hospitals are spending real money to collect revenue they were already owed.

  • Hospitals spent nearly $18 billion overturning denials in 2025, within roughly $43 billion spent overall pursuing payments insurers owe for care already delivered, per AHA estimates. Average denied MA claim value is around $1,000, up 22.4% year over year.

  • Underpayments are structurally larger and quieter. Medicare reimbursed hospitals at about 83 cents per dollar of cost in 2024, producing more than $100 billion in underpayments, and MA reimbursement fell 8.8% on a cost basis between 2019 and 2024, per AHA analysis.

  • AI works when the data underneath is connected. Black Book Research's 2025 evaluation found 83% of organizations reported AI-driven automation reduced claim denials by at least 10% within six months, and mature deployments reach 30% to 40% reductions. HFMA research found 67% of organizations expect AI and automation to drive the most significant impact on denials and underpayments over the next twelve months.

  • And it fails when the data is not. Industry analyses find roughly 4 of every 33 AI pilots reach production, an 88% failure rate at the prototype-to-production transition, with fragmented EHR integration and non-contextual data cited as the primary causes.

What is revenue leakage, and how is it different from denials?

Revenue leakage is the total gap between the revenue an organization earned by delivering care and the revenue it actually collects and keeps. Denials are one component. The larger components are quieter.

The full picture has five parts:

  1. Denials. A payer refuses the claim. Visible, worked, appealable, and measured. Roughly 11.65% of claims initially, with a meaningful share overturned when appealed.

  2. Underpayments. The payer pays, but less than the contract requires. Invisible in every standard workflow because the claim closes as paid. This is where the 1% to 11% of net patient revenue goes.

  3. Missed charge capture. Services delivered but never billed. Common in ancillary services, supplies, implants, and time-based procedures where documentation and charge entry are separated.

  4. Coding and documentation gaps. Care delivered at a higher acuity than the documentation supports, so the claim is billed accurately but at a lower level than the encounter warranted.

  5. Front-end preventable loss. Eligibility, authorization, and referral failures that convert a legitimate service into an unpaid one before the claim is ever created. Our patient access demo shows referral-to-first-visit funnel leakage on fictional data.

Denial management addresses category one. Most organizations have no systematic process for categories two through five, not because they do not care but because the detection requires cross-system context that no single application holds.

What are denial and underpayment rates in 2026?

Denials are up, and the amount at stake per denial is up faster than the denial count.

Initial denial rate: approximately 11.65%, with industry-wide estimates near 12% for 2026. That is more than one in nine claims requiring rework.

Denied amount growth: up 14% in hospital outpatient and 12% in inpatient year over year, meaning the mix of what gets denied is shifting toward higher-value claims.

Medicare Advantage: initial denial rates of 17% in the Health Affairs study, with 57% of denials overturned on appeal. MA denial rates rose 4.8% from 2023 to 2024. The average denied MA claim runs about $1,000, up 22.4% year over year.

Write-offs: 41% of providers absorb $5 million or more in annual denial write-offs, which is the portion of denied revenue that is never appealed or never recovered.

Accounts receivable pressure: true AR days rose 5.2% year over year in 2024, and AR over 90 days is averaging around 36% against a 15% to 20% benchmark.

Audit exposure: MDaudit's analysis across roughly 1.2 million providers and 4,500 facilities found a 30% year-over-year increase in at-risk audit amounts per customer, and an 18% increase in average at-risk amount per claim.

The 57% overturn rate deserves particular attention, because it defines the size of the opportunity. If more than half of MA denials are wrong, then the constraint is not whether you can win appeals. It is whether you can identify, prioritize, and work enough of them with the staff you have.

Why can't your existing RCM tools see underpayments?

Because detecting an underpayment requires comparing what was paid against what should have been paid, and "should have been paid" is a calculation nothing in the standard stack performs.

To compute expected reimbursement for a single claim you need:

  • The encounter detail: procedures, diagnoses, modifiers, units, place of service, provider, and payer plan, which live in the EHR and the practice management system.

  • The claim as submitted, which lives in the clearinghouse or billing system and may differ from what the EHR recorded after scrubbing.

  • The contract terms for that specific payer, plan, and effective date: fee schedule, percentage of Medicare, carve-outs, implant and drug pass-throughs, stop-loss and outlier provisions, and multiple-procedure reduction rules. These live in PDFs in a contracting folder, and often in a contract manager's memory.

  • The remittance detail: the 835, the adjustment codes, the payer's stated reason for each line-level difference.

  • The downstream adjustments: takebacks, offsets applied against unrelated claims, and recoupments that arrive months later and are often applied at the aggregate rather than claim level.

Practice management systems hold what was billed and what was paid. They do not hold a computable model of the contract, so they cannot say what should have been paid. That is why an underpaid claim posts as satisfied. The variance exists, but nothing in the system is capable of noticing it.

Contract management modules exist and help, but they typically model the primary fee schedule and not the carve-outs, pass-throughs, and reduction rules where the actual leakage concentrates. The complex provisions are exactly the ones payers apply inconsistently.

Why do healthcare AI RCM pilots fail after go-live?

Because the pilot ran on curated data and production runs on real data. Industry post-mortems consistently find that only about 4 of every 33 AI pilots reach production, an 88% failure rate at that transition, and the causes are structural rather than algorithmic.

The specific failure mechanisms:

The pilot excluded the hard cases. Pilot scope is controlled, the dataset is manually cleaned, and edge cases are excluded. Live operations supply nothing but edge cases. Denials cluster in exactly the populations a pilot trims: coordination of benefits, retroactive eligibility, out-of-network, secondary payers.

Data arrives fragmented and non-contextual. Clinical records, claims, scheduling, call center notes, and operational platforms all generate signal, and they are rarely unified. A model trained on a partial view of reality performs to the partiality of that view.

Multi-EHR reality. A system might run Epic in the hospital, eClinicalWorks in an affiliated clinic, and athenahealth in a specialty group. A model validated on Epic behaves differently against Cerner because the underlying data structures are not standardized. Any automation layer has to handle all of them, and most pilots handle one.

API constraints nobody planned for. Epic and Cerner throttle API requests, govern bulk extraction, and tightly control write-back permissions. An agent that needs continuous encounter context will hit those limits at production scale, having never approached them in a pilot.

No feedback loop. Tools that cannot retain outcome feedback, adapt to context, or improve over time plateau at their launch accuracy. In denial work, the outcome data (which appeals won, on what argument, with which payer) is the most valuable training signal available, and most deployments discard it.

Healthcare has invested more than $40 billion in AI while EMR fragmentation prevents much of it from delivering. The fix is not a better model.

What is a healthcare data layer?

A healthcare data layer is a governed layer between your clinical, financial, and operational systems and everything you use to make decisions, holding connections to those systems, a resolved identity model for patients, encounters, providers, payers, and contracts, and computable versions of the rules that determine what you should be paid.

Three components do the work.

Entity resolution. One patient exists as three MRNs across three sites. One provider has an NPI, several payer-specific IDs, and a name spelled two ways. One payer exists as a parent, several plans, and multiple network products with different fee schedules. Until these resolve to persistent identities, cross-site and cross-payer analysis is guesswork. Our multi-site consolidation demo shows fourteen clinics measured on one yardstick using fictional data.

Computable contracts. The contract PDF becomes structured, versioned, effective-dated logic: base fee schedule, carve-outs, pass-throughs, reduction rules, outlier provisions. This is the artifact that makes expected reimbursement calculable, and it is the single highest-value thing most organizations do not have.

Lineage and permissions. Every derived number traces to source records, and PHI access is enforced in the layer according to role and minimum necessary, not delegated to whatever tool is asking. In healthcare this is not optional architecture.

With those three in place, expected reimbursement becomes a computed field on every claim, and the variance between expected and actual becomes a work queue. We described the general pattern in What Is an AI Context Layer? and the graph mechanics in Knowledge Graphs for Enterprise AI.

Do you have to replace Epic, Cerner, or your practice management system?

No. Replacing an EHR to improve revenue cycle analytics would be one of the most expensive ways to solve a problem that does not require it. The EHR remains the clinical system of record and the PM system remains the billing system of record.

What changes is that the layer reads from them rather than asking them to do a job they were not built for. Practically that means respecting the constraints those systems impose: scheduled bulk extraction within governed limits rather than continuous real-time API calls, incremental syncs, and a local resolved model that agents query instead of hammering the EHR directly. Planning for API throttling in week two rather than discovering it in month six is one of the clearest differences between programs that reach production and programs that do not.

What systems does the layer need to connect?

Clinical. Epic, Cerner, Meditech, athenahealth, eClinicalWorks, Veradigm, or specialty-specific EHRs. Source of encounter detail, documentation, orders, and diagnoses.

Billing and claims. Practice management system, clearinghouse, 837 submissions and 835 remittances, denial and adjustment codes.

Contracts. Payer agreements, fee schedules, amendments, single-case agreements, letters of agreement. Usually PDFs, usually in a shared drive, occasionally only in email.

Scheduling and access. Referral management, prior authorization records, eligibility verification history. These determine preventable front-end loss.

Financial. General ledger, cost accounting, and supply or implant costs, which are required to move from revenue recovery to procedure-level profitability. Our procedure economics demo shows profit by surgeon and insurer on fictional data.

Operational. Time and attendance, OR block utilization, staffing. Needed to compute cost per case and contribution margin rather than revenue alone.

Payer correspondence. Denial letters, medical necessity requests, audit notices, appeal outcomes. This is where the argument that wins an appeal is recorded, and it is almost always unstructured.

How does underpayment detection actually work?

In five steps, executed per claim line rather than per claim.

  1. Compute expected reimbursement from the computable contract, at the line level, for the specific payer, plan, and date of service, applying carve-outs, pass-throughs, multiple-procedure reductions, and outlier provisions.

  2. Compare against the 835 remittance at line level. Aggregate-level comparison hides offsetting errors, which is exactly how systematic underpayment survives audit.

  3. Classify the variance. Contractual and correct, contractual and incorrect, patient responsibility, coding-driven, bundling or unbundling dispute, timely filing, or unexplained. Only a subset is recoverable, and separating them is what keeps the work queue credible.

  4. Prioritize by expected recovery value, which is the variance amount multiplied by the historical win rate for that payer, that variance type, and that argument. This is where the outcome feedback loop pays for itself: after two quarters you know which payers concede which arguments.

  5. Route with the argument pre-assembled. The staff member receives the claim, the contract clause, the remittance line, and the prior successful appeal language for that payer and variance type, rather than a variance number and a blank form.

The same machinery, run before submission rather than after payment, becomes denial prevention: flag the claim whose payer, procedure, and documentation pattern historically denies, and fix it while it is still cheap. Our revenue cycle recovery demo shows denials and underpayments with an appeal worklist on fictional data, and the full set of workflows sits on our healthcare page.

What can AI agents do here, and what should they not do?

They can do the reading and the assembly, at volume, continuously.

Appropriate:

  • Reading denial letters and payer correspondence to extract the actual reason, which frequently differs from the adjustment code.

  • Drafting appeals grounded in the specific contract clause and the documentation in the record, with citations.

  • Executing payer portal status checks and follow-up. Agentic execution of portal checks, appeals, and follow-up is now in production at scale, and 66% of revenue cycle leaders rated AI-powered denial follow-up as very important to their 2026 strategy.

  • Detecting variance patterns across payers that no human would find, such as a plan that systematically applies a reduction rule the contract does not permit.

  • Summarizing a case for a peer-to-peer review with the clinical facts a physician needs in one page.

Not appropriate without human sign-off: the coding decision, the medical necessity assertion, anything that constitutes a representation to a payer or regulator, and any number that reaches a board or lender without lineage. HFMA's finding that 67% of organizations expect AI to drive the largest impact on denials and underpayments is a statement about labor, not about judgment.

The results when the data layer exists are real: 83% of organizations in Black Book's 2025 evaluation reported at least a 10% denial reduction within six months, and mature deployments reach 30% to 40%. See AI Agents, Workflows, and Knowledge Graphs in Healthcare for the broader workflow landscape.

How do you handle PHI, HIPAA, and governance?

Decide it in week one, in writing, or the program stalls in month four. The non-negotiables:

  • A business associate agreement with any vendor touching PHI, and clarity on subprocessors including model providers.

  • Minimum necessary enforced in the layer. Role-based and row-level access, so a billing analyst sees claims and a clinical reviewer sees documentation, and neither inherits the other's scope because a dashboard was convenient.

  • No PHI in model training unless explicitly contracted and reviewed. Confirm retention and training policies in writing rather than in a sales conversation.

  • Full audit logging of who and what accessed which records, including agent actions. Agent activity is auditable activity.

  • De-identification for analytics where the analysis does not require identity, which is most portfolio-level and benchmarking work.

Data privacy concerns are consistently among the top reasons these programs stop. Handling them explicitly at the start converts a blocker into a checklist.

What does a realistic 90-day sequence look like?

Days 1 to 15: pick one payer and one service line. Not the whole book. Choose the payer with the largest denial or variance dollars, ideally an MA plan given the 17% initial denial rate and 57% overturn rate. Structure that payer's contract into computable form, including carve-outs. Get the BAA and access model signed.

Days 16 to 45: connect and compute. Connect the EHR, PM system, clearinghouse, and 835 feed for that service line. Resolve patients, providers, and payer plans. Compute expected reimbursement on the last twelve months of paid claims and produce the variance list.

Days 46 to 70: validate with the people who will work it. Sit with the revenue cycle team and adjudicate a sample of one to two hundred variances line by line. You are calibrating the classifier and earning the team's trust simultaneously. A variance queue the staff does not believe is a queue nobody works.

Days 71 to 90: work the queue and count the money. Submit appeals on the validated recoverable variances. Track submitted, won, lost, and dollars collected. This number is what funds the expansion to the next payer.

Then add payers one at a time, reusing the contract modeling pattern. Payer four takes a fraction of the time payer one took.

What goes wrong

Starting with all payers. Contract modeling is the expensive step and it is per payer. Doing twelve at once means finishing none before the budget review.

Comparing at claim level instead of line level. Line-level detection is where systematic underpayment shows up. Claim-level comparison nets out the errors you are hunting.

No feedback loop on appeal outcomes. Without capturing which arguments won against which payer, prioritization stays static and the second year performs no better than the first.

Treating it as an IT project. The contract knowledge lives with your contracting lead and the variance judgment lives with your revenue cycle veterans. If they are not in the room in week one, the model encodes assumptions that are wrong in ways only they would catch.

Measuring detected instead of collected. A dashboard showing $6 million detected is not $6 million. Report submitted, won, and collected.

How do you measure whether it worked?

  • Dollars recovered, net of cost to recover. The only number that matters at the board level.

  • Initial denial rate, against the roughly 11.65% industry baseline.

  • Underpayment recovery rate: identified recoverable variance versus collected.

  • Appeal win rate by payer and variance type, which should improve as the feedback loop matures.

  • Days in AR and percentage of AR over 90 days, against the 15% to 20% benchmark that most organizations currently miss at around 36%.

  • Cost to collect as a percentage of net patient revenue, which is where automation shows up if it is real.

Common questions about denial and underpayment detection

What is the average claim denial rate in 2026?

Approximately 11.65% initially, with industry-wide estimates near 12%. Medicare Advantage runs higher, at about 17% initial denials in a Health Affairs study covering roughly 30% of the MA market, with 57% of those denials overturned on appeal.

What is the difference between a denial and an underpayment?

A denial is a payer refusing to pay a claim, which creates a visible work queue item. An underpayment is a payer paying less than the contract requires while the claim posts as satisfied, so nothing surfaces it. Underpayments are estimated at 1% to 11% of net patient revenue annually.

How do you detect underpayments?

Compute expected reimbursement per claim line from a computable version of the payer contract, compare it against the 835 remittance at line level, classify the variance, prioritize by expected recovery value, and route it with the supporting contract clause attached. Claim-level comparison misses systematic underpayment because errors offset.

Do we need to replace our EHR to do this?

No. The EHR stays the clinical system of record. The layer reads from it on a governed schedule, respecting the API throttling and bulk extraction limits that Epic and Cerner enforce, and maintains its own resolved model for analysis.

Why did our AI denial pilot not scale?

Most commonly because the pilot ran on cleaned data with edge cases excluded, and production supplies mostly edge cases. Industry analyses find roughly 4 of 33 pilots reach production. Fragmented EHR integration, non-contextual data, unplanned API limits, and the absence of an outcome feedback loop are the recurring causes.

How much can AI actually reduce denials?

Black Book Research's 2025 evaluation found 83% of organizations reported at least a 10% denial reduction within six months of AI-driven automation, and mature deployments achieve 30% to 40% reductions. The variance between those outcomes is mostly a function of data readiness.

Is this worth doing for a physician group rather than a health system?

Often more so, because groups typically have fewer payers to model and no internal analytics team producing partial answers already. Modeling three commercial contracts and one MA plan covers most of the variance for many specialty groups.

How long until we see recovered dollars?

A focused single-payer, single-service-line program can be submitting validated appeals inside 90 days. Collection timing then follows the payer's appeal cycle, which is typically 30 to 90 days further out.

What is the most common mistake?

Buying denial analytics without computable contracts. Denial analytics tells you what was refused. Only a computable contract tells you what was underpaid, and the underpaid dollars are usually the larger pool.

Sources

  • Revecore, "Health System Denials and Underpayments Are Still Rising in 2026", which compiles the Health Affairs Medicare Advantage study (17% initial denial rate and 57% overturned on appeal, covering approximately 30% of the MA market), AHA figures (nearly $18 billion spent overturning denials in 2025, approximately $43 billion pursuing owed payments, Medicare reimbursement at 83 cents per dollar of cost in 2024 with over $100 billion in underpayments, MA reimbursement down 8.8% on a cost basis from 2019 to 2024), Becker's Hospital Review analysis (underpayments at 1% to 11% of net patient revenue, true AR days up 5.2%, AR over 90 days averaging 36% against a 15% to 20% benchmark), Fierce Healthcare (average denied MA claim approximately $1,000, up 22.4% year over year), and MDaudit (30% year-over-year increase in at-risk audit amounts across approximately 1.2 million providers and 4,500 facilities): revecore.com

  • Exactrx, "Hospital Denial Rate 2026" (41% of providers absorbing $5M+ in annual denial write-offs): exactrx.ai

  • Combine Health, "Top 10 AI Denial Management Solutions" reporting Black Book Research's 2025 AI RCM evaluation (83% of organizations reported at least 10% denial reduction within six months; mature deployments reaching 30% to 40%) and HFMA research (67% identifying AI and automation as the largest driver of impact on denials and underpayments over the next 12 months): combinehealth.ai

  • Mindbowser, "Why Healthcare AI Pilots Fail to Scale After Go-Live" (approximately 4 of 33 pilots reach production; fragmented EHR integration and non-contextual data as primary causes): mindbowser.com

  • Forbes, "The $40 Billion Healthcare AI Failure and the EMR Divide Sabotaging Progress": forbes.com

  • Robotics and Automation News, "Why EHR Integration is the Make-or-Break Factor in Healthcare AI Deployments" (Epic and Cerner API throttling, governed bulk extraction, and write-back permission constraints): roboticsandautomationnews.com

OutcomeCatalyst connects the systems you already run into a governed intelligence layer your team and your agents can reason over. Demos on this site use fictional data and are not clinical or billing advice. To see this on your own data, start a conversation.

Unified operating layer to harness artificial intelligence. Connect fragmented data, create agentic workflows, enable faster decisions across your company.

© 2026 OutcomeCatalyst. All rights reserved.