‹ Back to Blog

Data Strategy

Why AI Pilots Fail to Deliver ROI, and What the 5% Do Differently

Zach Shapiro

·

·

13 min

MPL Association, Medical Professional Liability Association

TL;DR: AI pilots fail to deliver ROI for a structural reason, not a technical one. Adoption is nearly universal and financial impact is rare. McKinsey's State of AI survey of about 2,000 respondents found more than 80% of organizations report no tangible enterprise-level EBIT impact from generative AI, while only 17% attribute 5% or more of EBIT to it. S&P Global Market Intelligence found the share of firms scrapping most of their AI initiatives rose to 42% in 2025 from 17% a year earlier. The differentiator is not model quality. McKinsey found high performers are nearly three times as likely to have fundamentally redesigned workflows, and Gartner found 63% of data-management leaders either lack AI-ready data practices or are unsure they have them.

Two facts about enterprise AI are both true right now, and they do not sit comfortably together. Adoption has effectively won: Deloitte's State of AI in the Enterprise, based on 3,235 board, C-suite and director-level leaders surveyed across 24 countries between August and September 2025, found only about a third of organizations are using AI at a surface level with little or no change to existing processes. The rest are redesigning processes or transforming around it.

And yet the money has not arrived. McKinsey's State of AI research found more than 80% of respondents say their organizations are not seeing a tangible impact on enterprise-level EBIT from generative AI. MIT Media Lab's Project NANDA, reviewing more than 300 enterprise deployments alongside 52 case studies and 153 leadership interviews, put the figure more bluntly: roughly 95% of generative AI pilots produced no measurable P&L return.

This piece is about the space between those two numbers. Not why AI fails in general, which is a tired question, but the specific mechanism that stops a working pilot from becoming a line on the income statement, and what the companies on the other side of that line did first.

Key takeaways

  • Adoption is not the constraint. Deloitte's survey of 3,235 leaders (August to September 2025, 24 countries) found roughly two thirds of organizations are redesigning processes around AI or transforming more deeply, not merely experimenting.

  • EBIT impact remains rare. McKinsey's State of AI research across about 2,000 respondents found more than 80% report no tangible enterprise-level EBIT impact from gen AI, and only 17% attribute 5% or more of EBIT to it.

  • Abandonment is accelerating. S&P Global Market Intelligence, surveying more than 1,000 respondents across North America and Europe, found 42% of businesses scrapped most of their AI initiatives in 2025, up from 17% the prior year, and that the average organization killed 46% of proof-of-concepts before production.

  • Agentic projects carry the same risk. Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.

  • The differentiator is workflow redesign. McKinsey found high performers are nearly three times as likely as others to have fundamentally redesigned individual workflows, and that reimagining workflows end to end correlates more strongly with value than any tooling choice.

  • The prerequisite is joined data. Gartner's survey of 248 data-management leaders found 63% either lack AI-ready data practices or are unsure whether they have them.

  • The systems are not connected. MuleSoft's 2026 Connectivity Benchmark reports organizations run an average of 897 applications, of which only 27% are connected to each other.

Why do AI pilots fail to deliver ROI?

Most AI pilots fail to deliver ROI because they automate a task inside a workflow rather than changing the workflow, and a faster task inside an unchanged process does not move an income statement. The saving is real and local. It is also absorbed by the steps on either side of it.

McKinsey's State of AI research found more than 80% of respondents report no tangible enterprise-level EBIT impact from generative AI, while 17% attribute 5% or more of their EBIT over the prior twelve months to it. That distribution is the important part. This is not a technology that works for nobody. It is a technology that works decisively for a minority and marginally for everyone else.

The mechanism is straightforward once you look at where the time goes. A drafting tool that writes a first-pass memo in two minutes instead of forty replaces the least expensive part of a diligence cycle. The expensive parts are assembling the inputs, resolving which of three systems holds the correct figure, and waiting for a person with authority to review. None of those moved, so the cycle time barely moved, so nothing reached the margin.

This is why the question a CFO should ask about any pilot is not whether it works. It is which step of which process it eliminates, and whether the steps around that one can absorb the change. For a fuller treatment of how that sequencing plays out by sector, see AI implementation by industry.

What "no EBIT impact" means at company scale

"No enterprise-level EBIT impact" does not mean the tools do nothing. It means the effect is too diffuse or too small to survive consolidation into a reported number. For a $10M to $1B company, that is a meaningful test.

McKinsey's finding that 17% of respondents attribute 5% or more of EBIT to gen AI is the more useful figure for an owner. Five percent of EBIT is a number a board notices. It implies the technology was applied to something structural rather than sprinkled across a department.

The gap between that 17% and the 80%-plus reporting nothing is not explained by budget. S&P Global Market Intelligence found cost, data privacy and security risks were the top obstacles cited by more than 1,000 respondents across North America and Europe, but cost obstacles do not explain why funded projects with working technology still produce no measurable return.

What explains it better is scope. A pilot scoped to a task produces task-level savings, which are real and unmeasurable at the enterprise level. A pilot scoped to a workflow produces cycle-time change, which shows up. The choice of scope is made before any model is selected, usually by someone deciding what seems safe to try first.

How many AI projects are abandoned before production?

Nearly half. S&P Global Market Intelligence found the average organization scrapped 46% of its AI proof-of-concepts before they reached production, and that the share of businesses abandoning most of their AI initiatives climbed to 42% in 2025 from 17% the year before.

That year-over-year move is worth sitting with. Abandonment rising from 17% to 42% in twelve months, across a survey of more than 1,000 respondents in North America and Europe, is not a story about the technology getting worse. It is a story about organizations reaching the point in the project where the integration work becomes visible, and deciding the remaining cost is not worth the modelled return.

Gartner projects a similar pattern for the next wave, predicting more than 40% of agentic AI projects will be canceled by the end of 2027 because of escalating costs, unclear business value and inadequate risk controls. Agents raise the stakes rather than changing the shape of the problem: an agent that cannot reach a system cannot act in it, so the integration bill arrives earlier and larger.

The practical read for an operator is that the pilot stage systematically underprices the project. What gets piloted is the model against a sample of clean data. What gets abandoned is the same model against the real data estate.

Why the model is almost never the bottleneck

The model is almost never the bottleneck because every serious model on the market can already read a document, extract a figure and draft a paragraph. The constraint is whether it can reach the documents, and whether the figures it extracts can be trusted against a system of record.

MIT's Project NANDA reached this conclusion directly. Reviewing more than 300 enterprise deployments, the researchers attributed the failure pattern not to model capability but to the gap between the model and the organization's workflows, structures and data. The technology performed. The integration did not exist.

Consider what a single useful question requires. Asking which customers are unprofitable after cost to serve means joining four systems. An ERP holds revenue and cost of goods. Carrier bills hold freight at shipment level. A spreadsheet holds rebate accruals, and a ledger account holds returns. Four systems, four owners, no shared key. The question is trivial. The plumbing is the entire job.

This is the part vendors skip in a demo, because a demo runs on prepared data. It is also why pilots impress and products disappoint. For the mechanics of how that joining actually works, see what an AI context layer is and the role of knowledge graphs in enterprise AI.

What do the companies with real returns do differently?

Companies with measurable AI returns redesign the workflow itself rather than accelerating a step inside it. McKinsey found high performers are nearly three times as likely as other organizations to say they have fundamentally redesigned individual workflows, and identified end-to-end workflow reimagining as one of the practices most strongly correlated with measurable value.

That finding is easy to nod at and hard to act on, because redesigning a workflow means changing who does what, in what order, with what authority. It is an operating decision, not a technology decision, and it is why the projects that work tend to be sponsored by a COO or an owner rather than by IT.

McKinsey also tested twelve adoption and scaling practices against EBIT impact and found positive correlations across all of them, with tracking well-defined KPIs for gen AI solutions carrying the strongest bottom-line effect. The pattern underneath both findings is the same: the organizations that got paid decided in advance what number they were trying to move, then rebuilt the process around moving it.

Deloitte's data points the same way. Among 3,235 leaders surveyed, roughly a third are redesigning key processes around AI and another third are transforming more deeply, while the remaining third are applying AI with little or no process change. That last third is where the pilots that never reach the P&L live.

Why AI-ready data is the prerequisite nobody budgets for

AI-ready data is the prerequisite most projects discover after approval rather than before it. Gartner's survey of 248 data-management leaders found 63% either lack AI-ready data practices or are unsure whether they have them, and Gartner has warned that AI projects will be abandoned through 2026 where those practices are not established.

"AI-ready" is a vague phrase, so it is worth making concrete. It means four things: the same entity resolves to one identity across systems, the fields carry the meaning your business assigns them rather than the meaning the vendor shipped, every figure keeps a link to the record it came from, and access rules follow the person the agent works for.

Miss the first and an agent double counts a customer that appears three ways. Miss the second and it reads a status code as a stage it is not. Miss the third and a committee refuses to act on the output, correctly, because nothing can be traced. Miss the fourth and you have a compliance problem rather than a productivity one.

None of these is exotic. All of them are unglamorous, they are invisible in a demo, and they are where the schedule goes. The honest version of an AI business case prices this work as the majority of the project, not as a footnote.

How disconnected is the average company, in numbers?

Badly, and more than most leaders assume. MuleSoft's 2026 Connectivity Benchmark reports that organizations run an average of 897 applications, and that only 27% of them are connected to one another.

Sit with the arithmetic. If fewer than three in ten applications are integrated, then the majority of the operational record of the business exists in systems that cannot answer a question jointly. Every cross-system question in that estate is currently answered by a person exporting, reconciling and re-keying, which is exactly the work that does not appear in any software budget.

This is the structural reason task-level automation disappoints. The task was never the constraint. The handoff between systems was, and a tool that speeds a task without touching the handoff leaves the cycle time roughly where it was.

It is also why integration cost is the single most underestimated line in an AI business case. The model is a commodity and the connectors are the project. For a fuller treatment of the trade-off between assembling that yourself and buying it, see buy versus build for an AI context layer.

Three ways companies try to close the gap

Most companies attempt one of three routes out of the pilot trap. They differ less in technology than in where the integration work lands and who carries the risk of it.

  • Point tools per department: What it costs Low per seat, high in aggregate, Time to value Weeks, Fails when Each tool solves one task, none share context, and the cross-system question is still unanswered

  • Build the layer in-house: What it costs Senior engineering time, ongoing, Time to value Quarters to years, Fails when Entity resolution and connector maintenance become a permanent team, competing with the product roadmap

  • Governed layer over existing systems: What it costs Implementation plus subscription, Time to value Weeks to months per workflow, Fails when The organization will not commit to a unit of value or a process owner, so nothing is redesigned

The third route is not automatically right. It is right when the constraint is that questions cross systems, which the MuleSoft integration figures suggest is the common case, and wrong when a single system genuinely holds the whole answer. A fuller comparison sits in AI company brain versus building your own RAG stack.

The failure mode nobody prices: the pilot that succeeds

The most expensive failure in enterprise AI is not the pilot that fails. It is the pilot that succeeds on a curated dataset, earns approval to scale, and then meets the real data estate. This is where budgets die, and it is almost entirely predictable in advance.

We have watched this pattern run the same way in every operation we have looked at. A pilot is scoped against a clean extract, often one region or one entity, prepared by the team that wants it to work. The results are genuine. Approval follows. Then scaling begins. The extract that took a week to prepare by hand must become a pipeline that runs nightly across every entity. That includes the two acquisitions still on their own systems, and the site that never migrated off the old ERP.

At that point three costs appear at once that no one modelled. Entity resolution stops being a spreadsheet lookup and becomes real infrastructure. Provenance becomes mandatory, because the moment output drives a decision, someone has to be able to trace a figure to its source. And access control stops being theoretical, because an agent reading across systems inherits every permission question the organization has been deferring.

S&P Global's finding that 46% of proof-of-concepts are scrapped before production is, in our reading, mostly this. Not projects that failed technically, but projects that were priced as a pilot and scoped as an integration programme. The tell is that abandonment rose to 42% in the same year agentic pilots proliferated, which is exactly what you would expect if the failure sits at the pilot-to-production boundary rather than in the models.

The way to defuse it is unglamorous: scope the pilot against the ugliest data you have, not the cleanest. If it works on the acquisition that never migrated, it will work everywhere. If it only works on the clean extract, you have learned that for the price of a pilot instead of the price of a programme.

How to size the EBITDA case before you commit

Size the case on cycle time and on work that is currently not done at all, because those are the two effects that survive consolidation into a reported number. Headcount savings are the weakest form of the argument and the hardest to collect.

Start with the unit your business counts: a claim, a policy, a property, a case, a part, a deal. Establish three figures. How long one unit takes end to end today. How many units never get worked because capacity ran out. What a worked unit is worth. Those three figures are usually available and rarely assembled.

The second effect matters more than most business cases admit. In every operation we have looked at there is a queue nobody works, not because the items lack value but because the hours ran out: referrals never contacted, submissions never quoted, claims never appealed, parts never compared across plants. Automation that expands throughput on that queue creates revenue rather than saving cost, and revenue is easier to defend in a board pack.

Then apply the discipline McKinsey's data points to. Name the KPI before the project starts, because tracking well-defined KPIs for gen AI showed the strongest correlation to bottom-line impact of the twelve practices tested. A project without a named number will produce a demonstration and no argument.

Finally, price the integration honestly, using the MuleSoft benchmark as a sanity check. If only 27% of the average estate is connected, assume the systems your workflow touches are not joined, and budget accordingly.

How OutcomeCatalyst fits

OutcomeCatalyst is a governed intelligence layer that connects the systems a company already runs into one structured, permissioned context that both people and AI agents can reason over, without replacing or migrating those systems.

On this specific problem, the order of work matters more than the tooling. We start from the unit of value your business counts and the number you are trying to move, not from an inventory of your software. We then resolve the entities that unit touches across the systems holding them. Every figure keeps a link to its source record, and the layer inherits your existing access rules rather than inventing new ones. Only then do we build the workflow on top, with a person approving anything that leaves the building.

The sequencing is the point. A workflow built before the layer exists is the pilot that succeeds and cannot scale. The layer built without a named unit of value is an infrastructure project with no sponsor. Both failure modes are common and both are avoidable.

You can see the shape of this in the staged workflows on this site, including claims and appeals in healthcare, direct spend and part consolidation in manufacturing, and underwriting and diligence in commercial real estate. Each is a worked example against realistic documents rather than a feature tour.

Common questions about AI pilots and ROI

Why do 95% of AI pilots fail?

MIT Media Lab's Project NANDA, reviewing more than 300 enterprise deployments, 52 case studies and 153 leadership interviews in 2025, found roughly 95% of generative AI pilots produced no measurable P&L return. The researchers attributed the pattern to organizations failing to integrate AI into workflows, structures and data rather than to model quality. The models performed. The surrounding process did not change.

What percentage of companies see real EBIT impact from AI?

McKinsey's State of AI research across about 2,000 respondents found more than 80% report no tangible enterprise-level EBIT impact from generative AI, while 17% attribute 5% or more of their EBIT over the prior twelve months to it. The distribution matters more than the average: the technology is producing decisive results for a minority and marginal results for most.

How long should an AI project take to show a return?

There is no defensible industry figure, and be sceptical of vendors who quote one. What the evidence supports is a sequencing rule: name the KPI before the project starts, because McKinsey found tracking well-defined KPIs carried the strongest correlation to bottom-line impact among twelve practices tested. A project without a named number cannot demonstrate a return regardless of how long it runs.

Is our data ready for AI?

Probably not, and that is the normal case rather than an indictment. Gartner's survey of 248 data-management leaders found 63% either lack AI-ready data practices or are unsure whether they have them. The practical test is narrower than a data-quality audit: can the same customer, patient or property be resolved to one identity across every system that holds it, and can any figure be traced to the record it came from.

Should we build the data layer ourselves or buy one?

It depends on whether integration is your business. Building means owning entity resolution and connector maintenance permanently, which becomes a standing team competing with your product roadmap. Buying means an implementation and a subscription, and the risk moves to vendor dependency. The deciding question is usually how many systems a single business question has to cross.

Why do agentic AI projects get canceled?

Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. Agents intensify the integration problem rather than changing it, because an agent that cannot reach a system cannot act in it, so the connector and permission work arrives earlier and at greater scale than with a drafting tool.

What should we automate first?

Start where a queue is going unworked for lack of hours rather than where a task is slow. Expanding throughput on work that is currently not done at all creates revenue rather than saving cost, and it avoids the trap of accelerating one step inside a process whose surrounding steps absorb the gain.

Sources

  • MIT Media Lab, Project NANDA, The GenAI Divide: State of AI in Business 2025, 300+ enterprise deployments, 52 case studies, 153 leadership interviews, August 2025 (95% of gen AI pilots produced no measurable P&L return; failure attributed to integration rather than model capability): mit.edu

  • McKinsey & Company, The State of AI: Global Survey, approximately 2,000 respondents, 2025 (more than 80% report no tangible enterprise-level EBIT impact; 17% attribute 5%+ of EBIT to gen AI; high performers nearly 3x as likely to have fundamentally redesigned workflows; KPI tracking most correlated with bottom-line impact): mckinsey.com

  • S&P Global Market Intelligence, AI adoption survey, more than 1,000 respondents across North America and Europe, 2025 (42% of businesses scrapped most AI initiatives, up from 17%; average organization abandoned 46% of proof-of-concepts before production; cost, privacy and security cited as top obstacles): ciodive.com

  • Gartner, press release, 25 June 2025 (more than 40% of agentic AI projects predicted to be canceled by end of 2027 due to escalating costs, unclear business value and inadequate risk controls): gartner.com

  • Gartner, data management survey, 248 data-management leaders, 2025 (63% either lack AI-ready data practices or are unsure whether they have them; AI projects at risk of abandonment through 2026 without AI-ready data): gartner.com

  • Deloitte, State of AI in the Enterprise, 3,235 board, C-suite and director-level leaders across 24 countries, surveyed August to September 2025 (roughly one third transforming deeply, one third redesigning key processes, one third applying AI with little or no process change): deloitte.com

  • MuleSoft (Salesforce), 2026 Connectivity Benchmark Report (organizations run an average of 897 applications, of which only 27% are connected to one another): mulesoft.com

  • Deloitte, press release on 2026 State of AI report, From Ambition to Activation (adoption maturity distribution and barriers to scaling): deloitte.com

OutcomeCatalyst connects the systems you already run into a governed intelligence layer your team and your agents can reason over. Demos on this site use fictional data. To see this on your own operation, start a conversation.

Unified operating layer to harness artificial intelligence. Connect fragmented data, create agentic workflows, enable faster decisions across your company.

© 2026 OutcomeCatalyst. All rights reserved.