‹ Back to Blog

Data Strategy

Custom AI vs Plug and Play: Why Purpose-Built Wins on ROI

Zach Shapiro

·

·

12 min

MPL Association, Medical Professional Liability Association

TL;DR: Custom AI vs plug and play is settled by the evidence, but not in the way either camp argues. MIT Media Lab's Project NANDA, reviewing more than 300 enterprise deployments, found AI sourced from specialized external vendors succeeded roughly 67% of the time against 33% for internally built tools, while general-purpose consumer tools scored high on satisfaction and reached production in only about 5% of evaluations. Purpose-built beats generic. It also beats do-it-yourself. The winning shape is a system fitted to one workflow and sourced from someone who has built it before.

The debate about custom AI versus plug and play is usually framed as a choice between speed and fit. Buy the generic tool and have something running on Monday, or build something bespoke and wait two quarters. Framed that way, plug and play looks obviously correct for a company that does not employ a machine learning team.

The data does not support that framing. MIT Media Lab's Project NANDA, drawing on more than 300 enterprise deployments, 52 case studies and 153 leadership interviews, found that AI systems sourced from specialized external vendors reached successful deployment about 67% of the time, roughly double the 33% success rate of tools built internally. Meanwhile general-purpose assistants that users genuinely liked were evaluated widely, reached pilot in about 20% of cases and production in about 5%.

Read together, those two findings say something more specific than "custom wins". They say fit to a workflow is what predicts success, that fit is usually bought rather than built, and that the tools which feel best to use are often the ones least likely to survive contact with a workflow that carries money or regulation.

Key takeaways

  • Specialized beats generic and beats DIY. MIT's Project NANDA found specialized external vendor tools succeeded in roughly 67% of deployments against about 33% for internally built systems, across more than 300 enterprise deployments reviewed in 2025.

  • Liking a tool is not the same as shipping it. The same research found general-purpose assistants were evaluated broadly but reached pilot in about 20% of cases and production in about 5%.

  • The failure is workflow fit, not model quality. MIT attributed the gap to learning, memory and workflow adaptation rather than to model capability.

  • Workflow redesign is the strongest correlate of value. McKinsey's State of AI research across about 2,000 respondents found high performers were nearly three times as likely to have fundamentally redesigned individual workflows.

  • Most estates cannot support a generic tool. MuleSoft's 2026 Connectivity Benchmark reports an average of 897 applications per organization with only 27% connected to one another.

  • Agentic projects fail on the same axis. Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 2027 on escalating costs, unclear business value and inadequate risk controls.

  • Readiness is the gating factor. Gartner's survey of 248 data-management leaders found 63% either lack AI-ready data practices or are unsure whether they have them.

What is the difference between custom AI and plug and play?

Custom AI is a system fitted to one workflow in one business, carrying that business's entities, rules and permissions. Plug and play is a general-purpose tool that any company can switch on, which by design knows nothing about your process until someone tells it.

The distinction is not about model choice, and it is worth being precise because vendors blur it. Both approaches typically run the same frontier models. What differs is everything around the model: whether it can reach your systems, whether it resolves your customer to one identity, whether it applies your underwriting criteria rather than a generic screen, and whether a figure it produces can be traced to a source record.

A plug-and-play assistant is excellent at tasks with no institutional context. Summarize this document. Draft this email. Rewrite this paragraph. These are genuinely valuable and they are also the tasks where the value is capped, because the work was never the constraint.

A purpose-built system is aimed at a workflow with institutional context: score this submission against our appetite, test this remittance against the contract we signed, rank these referrals by care at risk. Those questions cannot be answered by a tool that has not been connected to the systems holding the answer. For the underlying mechanics, see what an AI context layer is.

Why do plug and play AI tools demo well and stall in production?

Plug-and-play tools demo well because a demo runs on a document you hand the tool, and stall in production because production requires reaching systems nobody has connected. The demo tests the model. Production tests the integration.

MIT's Project NANDA quantified the drop. General-purpose tools were evaluated widely, reached pilot in roughly 20% of cases and reached production in about 5%, despite users rating them highly on immediate usability and satisfaction. That gap between satisfaction and deployment is the clearest signal in the report.

The mechanism is that a generic tool has no memory of your business between sessions and no path into your systems. It can read what you paste. It cannot read the rent roll in the data room, the 835 remittance in the clearinghouse or the shift notes in the maintenance log, which is where the answers to operational questions live.

MuleSoft's 2026 Connectivity Benchmark puts the scale of that problem plainly: an average of 897 applications per organization, of which only 27% are connected to one another. A tool that cannot cross that gap is limited to the work a person can already put in front of it, which is why the productivity gain is real, local and invisible at the P&L. This is the same pattern behind AI pilots that stall before EBIT.

What does the evidence say about custom versus bought AI?

The evidence favors specialized over generic and bought over built, in that order. MIT's Project NANDA found AI sourced from specialized external vendors succeeded in about 67% of deployments, roughly twice the approximately 33% success rate of internally developed tools.

That second comparison is the one most often skipped when this research gets quoted. "Custom" is frequently taken to mean "we will build it", and the data is unkind to that reading. Internal builds failed about two thirds of the time, which is worse than the specialized vendor route by a factor of two.

The reason is not talent. It is that the unglamorous parts of the problem, entity resolution across systems, connector maintenance as vendor APIs change, provenance on every field, and permission inheritance, are permanent work rather than project work. A team that builds them once has to keep them running forever, competing against the product roadmap that actually earns the company money.

The practical conclusion is a narrow one. Buy the fit, do not build the plumbing, and reserve internal engineering for the thing your customers pay you for. A fuller version of that trade-off sits in buy versus build for an AI context layer and in AI company brain versus building your own RAG stack.

Custom does not mean building it yourself

Custom AI means fitted to your workflow. It does not mean written by your engineers, and conflating the two is the most expensive misreading of the custom-versus-generic debate.

MIT's finding of roughly 33% success for internal builds against 67% for specialized vendors is the headline evidence, but the mechanism matters more than the ratio. When a company builds internally, it acquires four obligations at once: the connectors, the entity model, the governance, and the ongoing maintenance of all three as every upstream vendor ships changes.

Gartner's survey of 248 data-management leaders found 63% either lack AI-ready data practices or are unsure whether they have them. A company in that position is not one sprint away from a data layer. It is at the start of an infrastructure programme it did not budget for, which is precisely how a promising pilot turns into an abandoned one.

The middle path is what the successful cohort in the MIT data actually did. Buy a system that is purpose-built for the shape of your workflow, configure it against your own entities, rules and thresholds, and keep the customization at the level of business logic rather than plumbing. Configuration is cheap and reversible. Infrastructure is neither.

When is plug and play the right answer?

Plug and play is the right answer when the work is genuinely generic, when no institutional context is required, and when the output does not need to be traced. That covers more ground than purists admit.

Drafting, summarizing, translating, rewriting and first-pass research are all real work, and a general-purpose assistant handles them without any integration at all. Deloitte's State of AI in the Enterprise, surveying 3,235 leaders across 24 countries between August and September 2025, found worker access to AI tools rising sharply, and that access produces genuine individual productivity.

The mistake is not using plug-and-play tools. The mistake is expecting them to move a number on the P&L. Individual time savings distributed across a department rarely survive consolidation, which is consistent with McKinsey's finding that more than 80% of organizations report no tangible enterprise-level EBIT impact from generative AI.

A reasonable stack for a $10M to $1B company runs both. General-purpose assistants for individual knowledge work, and a purpose-built system for the two or three workflows that carry the money. The error is buying only the first and calling it an AI strategy.

Which workflows justify a purpose-built system?

Workflows justify a purpose-built system when they cross systems, carry material money or regulatory exposure, repeat at volume, and currently run below capacity. All four conditions matter, and the fourth is the one most often missed.

Crossing systems is the technical trigger. If answering the question requires joining an ERP to a carrier bill to a credit memo, no generic tool can reach it, and the integration is the project regardless of which model you choose.

Material exposure is the governance trigger. Where an output drives a payment, a filing, a clinical decision or a bid, provenance stops being nice and becomes mandatory, because someone will eventually be asked to defend the figure. Generic tools do not carry provenance because they were never connected to a source record.

Volume and unworked capacity are the economic triggers. The strongest business cases we see are not about doing existing work faster. They are about work that is currently not done at all: submissions never quoted, claims never appealed, referrals never contacted, parts never compared across plants. Expanding throughput on an unworked queue creates revenue rather than trimming cost, which is a far easier argument to carry to a board. You can see that pattern in the staged workflows for insurance submission and triage and healthcare referral and patient access.

Custom AI vs plug and play, compared

The three routes differ less in model quality than in where the fit work lands and who maintains it over time.

  • Plug and play assistant: What it costs Low per seat, Time to value Days, Fails when The question crosses systems, or an output must be traced to a source record

  • Build custom in-house: What it costs Senior engineering, permanently, Time to value Quarters to years, Fails when Connector and entity maintenance becomes a standing team competing with the product roadmap

  • Purpose-built system on your context layer: What it costs Implementation plus subscription, Time to value Weeks to months per workflow, Fails when The business will not name a unit of value or a process owner, so nothing gets redesigned

None of these is universally correct. The first is right for generic individual work, the third for workflows that cross systems and carry money, and the second mainly when integration itself is the product you sell.

The failure mode: buying a generic tool for a regulated workflow

The most common expensive mistake we watch companies make is not choosing wrong between custom and generic. It is choosing generic for a workflow that will later be audited, then discovering at audit that nothing can be traced.

The sequence is consistent. A team adopts a general-purpose assistant, which is genuinely useful. Someone begins using it for something load-bearing: summarizing a policy file, drafting an appeal, checking a contract term. It works well enough that the practice spreads informally. No procurement decision is ever made, because no procurement decision was required.

Then the workflow is examined. A regulator asks for the basis of a determination, an investor asks why a valuation moved, or a payer disputes an appeal. The answer has to be a source document and a line item, and what exists is a chat transcript. The output was probably correct. It is simply not defensible, and undefensible is functionally the same as wrong in a regulated process.

The second half of the failure is quieter. Because the tool never connected to a system of record, none of the work it did accumulated. Every session started from nothing. A purpose-built system on a context layer gets better as the entities resolve and the corpus grows; a generic assistant is exactly as capable in month twelve as in week one. MIT's framing of the gap as one of learning, memory and workflow adaptation is describing this directly.

The practical defense is to decide in advance which workflows are load-bearing, and to keep general-purpose tools away from them. Not because the tools are bad, but because they were built for a different job.

How to decide, in one test

Ask what has to be true for the output to be defensible, and let the answer choose the tool. If the output needs to be traced to a source record, you need a purpose-built system on connected data. If it does not, plug and play is sufficient and cheaper.

Run the test on a real example rather than a category. Take one output your business produces weekly, and ask three questions. Which systems hold the inputs. Who would have to defend the figure if challenged, and with what evidence. What happens if the number is wrong.

If the inputs sit in one system, nobody defends anything and a wrong number is embarrassing rather than costly, a general-purpose tool will serve. If the inputs cross three systems, a committee or a regulator can demand the basis, and a wrong number moves money, no amount of prompt engineering closes that gap.

Then apply McKinsey's discipline. Their research across roughly 2,000 respondents found tracking well-defined KPIs for gen AI solutions carried the strongest correlation to bottom-line impact among twelve practices tested. Name the number before you choose the tool, and the choice usually makes itself. For sector-specific starting points, see AI implementation by industry.

How OutcomeCatalyst fits

OutcomeCatalyst is a governed intelligence layer that connects the systems a company already runs into one structured, permissioned context that people and AI agents can both reason over, without replacing or migrating those systems.

The position on this question is deliberately narrow. We do not think most companies should build a data layer, and we do not think a general-purpose assistant will move an operating number. What works is a system purpose-built for the shape of a workflow, configured against your entities and thresholds rather than rewritten in code, with provenance on every field and your existing access rules inherited rather than reinvented.

In practice that means starting from the unit your business counts, resolving the entities that unit touches across the systems that hold them, and only then building the workflow on top with a person approving anything that leaves the building. The customization that matters is business logic, which you should own and change freely. The plumbing underneath it is not where a company of this size should be spending senior engineering time.

Worked examples of that shape live in manufacturing direct spend and commercial real estate underwriting and diligence, each built against realistic documents rather than a feature tour.

Common questions about custom AI vs plug and play

Is custom AI better than off-the-shelf AI?

For workflows that cross systems or carry regulatory exposure, yes. MIT's Project NANDA found specialized vendor systems succeeded in roughly 67% of deployments against about 33% for internal builds, while general-purpose tools reached production in only about 5% of evaluations despite high user satisfaction. For generic individual work such as drafting and summarizing, off-the-shelf is sufficient and cheaper.

Does custom AI mean we have to build it ourselves?

No, and the data argues against it. Internal builds succeeded about half as often as specialized vendor systems in MIT's review of more than 300 deployments. The durable position is buying a system purpose-built for your workflow shape and customizing it at the level of business rules, rather than owning connectors, entity resolution and governance permanently.

How much does a purpose-built AI system cost compared with a per-seat tool?

Per-seat tools are cheaper per user and the comparison misleads, because they solve a different problem. The honest comparison is between a purpose-built system and the fully loaded cost of the manual process it replaces, including the work currently not being done for lack of capacity. Ask any vendor to price against that baseline rather than against a seat licence.

Can we start with plug and play and move to custom later?

Yes, and that is a reasonable sequence, provided you do not let general-purpose tools become load-bearing in a workflow that will later be audited. The migration cost is not technical. It is that informal practice hardens into process, and unwinding it is a change-management problem rather than a software one.

Why do generic AI copilots underperform on operational work?

Generic copilots have no path into the systems that hold operational data and no memory of your business between sessions. MuleSoft's 2026 Connectivity Benchmark reports organizations run an average of 897 applications with only 27% connected, so most operational questions require joining systems the copilot cannot reach.

What makes an AI system purpose-built rather than just configured?

Purpose-built means the entity model matches your business, the rules encode your actual thresholds, every figure carries provenance to a source record, and access rules follow the person the agent works for. Configuration alone changes prompts and settings. Purpose-built changes what the system knows and what it can prove.

Should a $10M to $1B company use both?

Usually yes. General-purpose assistants for individual knowledge work, and one purpose-built system for the two or three workflows that carry the money. The failure pattern is buying only the first, distributing modest individual time savings across departments, and finding nothing has changed at the operating line.

Sources

  • MIT Media Lab, Project NANDA, The GenAI Divide: State of AI in Business 2025, 300+ enterprise deployments, 52 case studies, 153 leadership interviews, August 2025 (specialized external vendor systems succeeded in ~67% of deployments against ~33% for internal builds; general-purpose tools reached pilot in ~20% and production in ~5% of evaluations; failure attributed to learning, memory and workflow adaptation): mlq.ai

  • McKinsey & Company, The State of AI: Global Survey, approximately 2,000 respondents, 2025 (more than 80% report no tangible enterprise-level EBIT impact; high performers nearly 3x as likely to have fundamentally redesigned workflows; KPI tracking most correlated with bottom-line impact of twelve practices tested): mckinsey.com

  • MuleSoft (Salesforce), 2026 Connectivity Benchmark Report (organizations run an average of 897 applications, of which only 27% are connected to one another): mulesoft.com

  • Gartner, press release, 25 June 2025 (more than 40% of agentic AI projects predicted to be canceled by end of 2027 due to escalating costs, unclear business value and inadequate risk controls): gartner.com

  • Gartner, data management survey, 248 data-management leaders, 2025 (63% either lack AI-ready data practices or are unsure whether they have them): gartner.com

  • Deloitte, State of AI in the Enterprise, 3,235 board, C-suite and director-level leaders across 24 countries, surveyed August to September 2025 (adoption maturity distribution; roughly one third applying AI with little or no process change): deloitte.com

  • S&P Global Market Intelligence, AI adoption survey, more than 1,000 respondents across North America and Europe, 2025 (42% of businesses scrapped most AI initiatives, up from 17% the prior year; average organization abandoned 46% of proof-of-concepts before production): ciodive.com

  • Fortune, coverage of MIT Project NANDA findings, 18 August 2025 (independent reporting of the 95% no-return finding and the buy-versus-build success gap): fortune.com

OutcomeCatalyst connects the systems you already run into a governed intelligence layer your team and your agents can reason over. Demos on this site use fictional data. To see this on your own operation, start a conversation.

Unified operating layer to harness artificial intelligence. Connect fragmented data, create agentic workflows, enable faster decisions across your company.

© 2026 OutcomeCatalyst. All rights reserved.