OutcomeCatalyst

The BrainHow our platform trains your agents

Book a strategy call

Data Strategy

18 min read

Context Engineering for Enterprise AI Agents: A Build Guide (2026)

Context Engineering for Enterprise AI Agents: A Build Guide (2026)

Context Engineering for Enterprise AI Agents: A Build Guide (2026)

Context engineering vs prompt engineering, the eight parts of a context layer, MCP access to company data, evals, and a 90-day build plan.

Network cables neatly connected in a server rack

Zach Shapiro

TL;DR: Context engineering is the discipline of deciding what information, definitions, permissions, memory and tools an AI agent sees at each step, so it can reason over your business instead of guessing. Prompt engineering tunes the instruction; context engineering builds the system that fills the context window with the right facts, from the right systems, under the right access rules. For enterprise agents that means a context layer: unified data with resolved entities, shared metric definitions, governance, memory, MCP tool access, citations and freshness signals. Build it in phases, one workflow and one countable unit at a time, and measure it with evals and traces before you scale.

An insurance operations lead asks an agent a simple question: "Which of our accounts renewing in Q1 had a loss ratio above 65% last year?" The agent answers in seconds. Eleven of the fourteen accounts are wrong. The policy admin system calls the insured "Harbor Point Holdings LLC," the claims system calls it "Harborpoint Hldgs," the CRM has two records for it, and "loss ratio" in finance includes IBNR while the underwriting team's version does not. The context was broken, not the model.

That gap is what context engineering exists to close. The term took off in mid 2025, when Andrej Karpathy described it as "the delicate art and science of filling the context window with just the right information for the next step," and Anthropic followed with an engineering guide. Most of what ranks for the phrase is written for developers building coding agents. This post is for CTOs, heads of data and AI, data engineers and technical operators in CRE, insurance, healthcare and manufacturing who have to make it work against an ERP, a CRM, an industry system of record and a decade of PDFs.

If you want the plain-language definition of the layer itself, start with our explainer on what an AI context layer is. If you want the data plumbing in depth, read data unification for an intelligence layer. This post treats context engineering as a discipline: components you can assign to owners, and a phased build plan with evals attached.

Key takeaways

  • Context, not model choice, is the usual failure point. MIT NANDA's 2025 "GenAI Divide" report attributes the high rate of stalled enterprise pilots to brittle workflows and weak contextual learning rather than model quality.

  • More context is not better context. Chroma's 2025 "Context Rot" study tested 18 LLMs and found performance "grows increasingly unreliable as input length grows."

  • Data readiness decides survival. Gartner predicts that through 2026 organizations will abandon 60% of AI projects unsupported by AI-ready data, and found 63% of organizations lack or are unsure they have the right data management practices for AI.

  • MCP is now the default plumbing for tool access. Anthropic reported more than 10,000 active public MCP servers when it donated the protocol to the Linux Foundation's Agentic AI Foundation in December 2025.

  • Quality, not cost, blocks production. In LangChain's State of Agent Engineering survey of 1,340 practitioners, 32% named quality as the top barrier, while 89% had observability and only about 52% ran offline evals.

  • Name the countable unit first. A context layer without a target metric (minutes per submission, days in A/R, deals screened per week) cannot show ROI.

What is context engineering?

Context engineering is the practice of designing everything an AI model sees at inference time, not just the prompt: retrieved data, entity records, business definitions, prior decisions, tool outputs, memory and the rules about who may see what. Anthropic defines it as "the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference," including information that lands in the window outside the prompt itself.

The useful mental model is a budget. Anthropic's guide calls it an "attention budget" and says context "must be treated as a finite resource with diminishing marginal returns." So the question is not "how do we give the agent everything?" It is "what is the smallest set of high-signal facts this step needs, fetched with permission and with proof of origin?"

In a consumer chatbot, that problem is mostly about conversation history. In an enterprise, it is mostly about data. The facts an underwriting agent needs live in Guidewire or Duck Creek, in loss runs that arrive as scanned PDFs, in Applied Epic, and in the head of a senior underwriter. Enterprise context engineering is half retrieval design and half data architecture.

Context engineering vs prompt engineering vs RAG: what is the difference?

Prompt engineering shapes how the model behaves; context engineering decides what the model knows and can do at each step; retrieval augmented generation (RAG) is one technique inside context engineering for pulling documents into the window. They are nested, not competing.

  • Prompt engineering works on the instruction: role, tone, output format, examples, reasoning steps. It matters, but it cannot fix a wrong number coming out of the data.

  • RAG retrieves relevant chunks of text (usually from a vector index) and inserts them into the prompt. It is strong on policy manuals and lease clauses, weak on joins, aggregations and definitions.

  • Context engineering is the system design around both: sources, unification, canonical definitions, memory, tools, citations and freshness.

Here is an arguable position, and we hold it: for operational agents in regulated industries, vector RAG over raw documents should be the last retrieval path you build, not the first. Most operator questions ("which tenants are 60 days past due," "which claims were underpaid against contract") are structured questions about entities. Answer those from a resolved, governed data model, and use document retrieval for the clauses that explain the numbers. Teams that start with a vector index over the shared drive ship a fast demo, then spend months explaining why the agent's numbers do not match the GL.

What goes into a context layer for AI agents?

A context layer is the governed system that supplies agents with trustworthy facts, definitions, memory and tools. It has eight components. Stalled agent projects are usually missing several of them.

1. Data unification

Agents need one queryable view across the systems that run the business: the ERP (NetSuite, SAP, Epicor), the CRM (Salesforce, HubSpot), the industry system of record (Yardi or MRI for property, Guidewire or Duck Creek for policy, Epic or athenahealth for clinical and revenue cycle, JobBOSS or ProShop for job shops), and the documents that never made it into any of them. Unification means a shared model of the business that each source maps into, with lineage back to the original record.

2. Entity resolution and master data

Entity resolution is the process of deciding that "Harbor Point Holdings LLC," "Harborpoint Hldgs" and CRM account 0034411 are the same insured. Without it, every aggregate an agent computes is quietly wrong. In CRE it is one tenant with three legal entities across a guarantor structure; in healthcare, payers whose plan names differ between the 835 and the contract; in manufacturing, one part with a customer number, an internal number and a SolidWorks revision the ERP never picked up.

Good entity resolution combines deterministic keys (tax ID, NPI, APN, part number) with probabilistic matching on names and addresses, a confidence score, and a human review queue for the gray zone. The output is a stable ID the agent can trust and a match record it can cite.

3. Ontology and metric definitions

An ontology says what things exist (property, lease, tenant, policy, claim, encounter, work order) and how they relate. Metric definitions say how you compute the numbers the business argues about: NOI, loss ratio, net collection rate, OEE, cost to serve. This is where the semantic layer lives. A common failure looks like this: the agent writes valid SQL that computes the wrong metric, for example summing a column that includes canceled orders. A governed definition that the agent must call, instead of re-deriving, removes that whole class of error. More on how ontology and metric layers compare in our piece on Databricks Genie, ontologies and the context layer.

4. Permissions and governance

The agent should never see more than the human it acts for. That means row and column level security inherited from source systems, PHI and PII handling, purpose limits, and an audit log of every read and write. Our own platform is HIPAA-aligned and SOC 2 Type 2 aligned, with the formal audit underway and expected to complete before year-end, and permission checks happen in the context layer, not the prompt.

5. Memory

Enterprise memory has three tiers. Working memory is the task scratchpad. Episodic memory is what happened on previous runs: which exceptions were approved, which broker was asked for a corrected schedule last month. Institutional memory is the decision history and rules of thumb your veterans carry, for example "we never accept a loss run older than 90 days" or "this payer always downcodes 99214 on the first pass." Anthropic's guide names structured note-taking as a core long-horizon technique; in an enterprise, those notes need an owner, a retention policy and a way for humans to correct them.

6. Tool access via MCP

Tools are how agents read live state and take action: query a rent roll, pull a claim, create a draft endorsement, open a purchase order. The Model Context Protocol gives you a standard way to expose those tools. More on MCP design below.

7. Provenance and citations

Every fact an agent uses should carry a pointer back to its source: the Yardi lease ID, the page of the loss run, the 835 line, the work order. Citations let a reviewer approve in thirty seconds instead of redoing the work. They are also your answer when an auditor asks how a decision was made.

8. Freshness

Different systems run on different clocks. The GL closes monthly, the PMS updates nightly, claims update in near real time. Stamp every fact with an "as of" time. An agent that compares a March rent roll to an April AR aging without saying so will produce a confident and wrong delinquency report.

How do you give AI agents access to company data with MCP?

You expose a small number of well-described, permission-aware tools through MCP servers that sit on top of your context layer, not directly on top of raw source databases. The agent calls the tool; the tool enforces identity, applies canonical definitions, and returns a compact, cited result.

MCP has become the default interface. When Anthropic donated it to the Linux Foundation's Agentic AI Foundation in December 2025, it reported more than 10,000 active public MCP servers and adoption in ChatGPT, Cursor, Gemini, Microsoft Copilot and VS Code. Standardization also makes it trivial to point an agent at a raw Postgres replica of your ERP, which is one of the most common enterprise MCP mistakes.

MCP server design patterns for enterprise data

  • Business tools, not tables. Expose get_tenant_delinquency(property_id, as_of) rather than run_sql. A SQL tool invites the agent to invent a definition.

  • Identity passthrough. The MCP server should act as the requesting user (OAuth on-behalf-of or equivalent), so source-system permissions apply.

  • Small, structured outputs. Return the twenty rows that matter with IDs and an "as of" stamp, not a 4 MB JSON dump. Anthropic's guide recommends just-in-time retrieval for this reason.

  • Read and write separated. Keep read tools broad and write tools narrow, with drafts by default and an explicit human approval step for anything that changes a system of record.

An example: CRE underwriting

An acquisitions analyst asks an agent to screen an offering memorandum. The agent extracts the rent roll from the OM, pulls comps from licensed CoStar data, and checks how similar assets in the firm's own Yardi and ARGUS files performed against pro forma. The context layer resolves OM tenants against existing exposure, flags overlap, and cites every number. That is the pattern behind our underwriting and diligence agent, and it only works because the tenant entities were resolved before the agent ever ran.

How do you unify ERP, CRM, industry systems and documents so agents can act on them?

Map every source into a shared business model with stable IDs, keep lineage to the original records, and extract structured fields from documents into the same model rather than leaving them in a separate vector store.

In practice the work runs in five passes.

  1. Inventory. List the systems, owners, refresh cadence, and access method for each (API, replica, SFTP, or "Jen emails it on Fridays"). Include the spreadsheets.

  2. Model. Define the core entities and relationships for the target workflow only. Do not model the whole enterprise up front.

  3. Resolve. Run entity resolution across sources and publish a crosswalk with confidence scores. Route low-confidence matches to the people who know the data.

  4. Extract. Pull structured fields from documents (loss runs, ACORD forms, leases, EOBs, purchase orders, certs of insurance) into the same model, with page-level citations. Keep the full text indexed for clause-level retrieval.

  5. Connect. Store the resolved entities and relationships in a form agents can traverse. A knowledge graph is a natural fit when questions hop across relationships (tenant to guarantor to parent to other leases), which is why we build on one.

Vertical examples

  • Commercial real estate. Yardi or MRI holds leases and receivables, ARGUS holds the underwriting model, and amendments live in PDFs. When an amendment updates Yardi but not ARGUS, the two disagree on lease terms. Reconcile before any reporting agent drafts an investor letter.

  • Insurance. Submissions arrive as ACORD forms, SOVs in Excel and loss runs in a dozen carrier formats, while policy admin sits in Guidewire or Duck Creek. A triage agent needs the insured resolved across all of them to know whether this is a re-marketed account you declined last year.

  • Healthcare. Underpayment detection means joining an 835 remittance line from the clearinghouse to the contracted rate for that payer, plan and CPT code, with encounters in Epic or athenahealth. Without resolved payers and plans, the join fails.

  • Manufacturing. A quote agent needs the part, revision and routing resolved across the ERP, the shop floor system and SolidWorks PDM, or it will quote against last year's cycle time.

How do you evaluate and observe a context-engineered agent?

Evaluate the context and the answer separately: test whether the right facts reached the model (retrieval and resolution quality), then whether the model used them correctly (answer quality), and trace every production run.

LangChain's 2026 State of Agent Engineering survey found that 89% of respondents had observability on their agents but only 52.4% ran offline evaluations on test sets. Traces tell you what happened; evals tell you whether it was right before a customer finds out.

What to evaluate

  • Entity resolution precision and recall on a labeled sample.

  • Metric fidelity: does the agent's loss ratio, NOI or net collection rate match the number finance signs off on?

  • Retrieval relevance: for each question in the golden set, did the right lease clause or claim line make it into context?

  • Citation validity: does every cited source exist and actually support the claim?

  • Permission tests: red-team prompts that try to pull data the requesting user should not see.

  • Freshness handling: cases with deliberately mismatched as-of dates, where the right answer is to flag the gap.

What to observe in production

Log every tool call, the tokens returned, the sources cited, the final output, and the human reviewer's decision (accepted, edited, rejected, with a reason code). The reviewer's edit is the most valuable signal you have. It becomes tomorrow's eval case.

What should you measure to prove ROI?

Pick one countable unit per workflow, baseline it before you build, and report the change in that unit, not "hours saved" estimates. Gartner's warning that agent projects are being canceled for "unclear business value" is mostly a warning about teams that skipped this step.

Good countable units are boring and specific: minutes per submission triaged, days in A/R, denials overturned per FTE, deals screened per analyst per week.

A worked example (hypothetical numbers)

A hypothetical specialty insurer gets 400 submissions a month, and assistants spend 55 minutes on each clearing, keying ACORD data and pulling history. Fully loaded cost is $60 per hour.

  • Baseline: 400 x 55 minutes = 22,000 minutes, or about 367 hours a month, roughly $22,000 a month in handling cost.

  • After a triage agent on a context layer with resolved insureds and extracted loss runs, assume handling drops to 20 minutes per submission, with a human reviewing every draft.

  • New load: 400 x 20 = 8,000 minutes, about 133 hours, roughly $8,000 a month. Savings: about $14,000 a month, or $168,000 a year.

  • The revenue side is often larger: if faster turnaround means fewer submissions age out unquoted, track quote rate and time to quote alongside minutes per submission.

These numbers are illustrative only. If you cannot write this math down before you build, you are not ready to automate the workflow. More on the gap between delivered value and realized ROI in why AI pilots stall before EBIT. Built is not adopted, and adoption is what the P&L sees.

What should agents not do? Where humans stay in the loop

Agents should do the reading, reconciling and first-draft work; humans should keep the judgment calls, the commitments, and anything that is hard to reverse. Citations make the review fast; they do not remove it.

  • Binding risk or declining it. An agent drafts the recommendation; an underwriter decides.

  • Clinical decisions and anything touching patient care. Revenue cycle agents can prepare an appeal; they should not decide medical necessity.

  • Price commitments to customers. A quote agent drafts; an estimator or sales lead releases.

  • Writes to systems of record without an approval step, until evals and reviewer data show a specific write type is safe.

  • Low-confidence entity matches. Route them to a person. One wrong merge contaminates thousands of answers.

There is also a category of work you should not automate yet: workflows with no agreed definition. If finance and operations still argue about how to compute cost to serve, an agent will only automate the argument. Settle the definition, encode it in the context layer, then automate.

How do you build a context layer? A phased, 90-day plan

Start with one workflow, one countable unit, and the minimum set of sources it needs; prove it with evals and human review; then widen the layer to the next workflow, reusing the entities and definitions you already resolved.

Days 1 to 15: choose and baseline

  1. Pick one workflow where the labor is reading and reconciling, the volume is real, and a senior person can judge output quality quickly. Submission triage, denial appeals and OM screening are common picks.

  2. Name the countable unit and baseline it from system data, not interviews.

  3. Inventory the sources the workflow touches and the definitions it depends on.

  4. Build a golden set of 50 to 100 real historical cases with the correct answers.

Days 16 to 45: build the minimum context layer

  1. Connect the sources with read access and lineage. Start with nightly loads; real time can wait.

  2. Model the core entities for this workflow and run entity resolution. Publish the crosswalk and review the low-confidence queue.

  3. Encode the two to five metric definitions the workflow depends on, signed off by their owners.

  4. Extract structured fields from the relevant document types with page citations.

  5. Wrap the layer in a handful of MCP tools with identity passthrough and as-of stamps.

  6. Run the golden set. Fix the context, not the prompt, until metric fidelity and citation validity are where the reviewers need them.

Days 46 to 75: shadow mode

  1. Run the agent beside the team on live work. Humans do the job as usual and grade the agent's draft.

  2. Capture every edit with a reason code and feed it into the eval set.

  3. Watch tokens per step, tool error rates, and freshness flags in your traces.

Days 76 to 90: assisted production and the next workflow

  1. Switch to agent first draft with human approval on the live queue. Report the countable unit weekly against baseline.

  2. Pick the adjacent workflow that reuses the most resolved entities. If the second workflow is not faster than the first, your layer is too specific.

  3. Make the build vs buy call with real data in hand. Some teams with strong data engineering should build. The most expensive answer is no decision: another quarter of pilots on unreconciled data.

What does enterprise AI architecture look like in 2026?

A durable 2026 enterprise AI architecture has four layers: systems of record at the bottom, a governed context layer above them, an agent and orchestration layer that reaches the context layer through MCP tools, and a human review layer where people approve, edit and teach. Models are swappable parts within the agent layer.

The design choice that matters most is where business meaning lives. If definitions, entity IDs and permissions live inside each agent's prompts, you end up with ten agents that disagree about the same number. If they live in the context layer, every agent and dashboard shares them. That is the case for treating context as infrastructure, not prompt writing.

Two habits keep each step small, which is what the context rot research argues for: scoped sub-agents that hand back condensed summaries, and compaction of finished steps into cited notes instead of carrying raw tool output forward.

This is the architecture behind the OutcomeCatalyst company brain: one governed layer over the systems you already run, with agents that reason over it the way your veterans would, and people who stay in charge of the calls that matter.

Frequently asked questions

What is context engineering in simple words?

It is the work of making sure an AI agent sees the right facts, definitions, history and tools at each step, and nothing it should not see. Most of it is data and system design.

Is context engineering the same as RAG?

No. RAG is one retrieval technique that pulls documents into the prompt. Context engineering also covers structured data, entities, definitions, permissions, memory, tools, citations and freshness.

How do I connect an AI agent to my ERP or CRM safely?

Put an MCP server in front of a governed layer, not the raw database. Expose narrow business tools, pass the user's identity through so source permissions apply, keep writes behind human approval, and log every call.

What is context rot and how do I avoid it?

Context rot is the drop in model accuracy as input length grows, documented across 18 models in Chroma's 2025 research. Avoid it by retrieving just in time, returning compact structured results, compacting finished steps into notes, and splitting long tasks across scoped sub-agents.

Who should own context engineering inside a company?

The head of data or data platform usually owns the layer, with metric definitions owned by the business teams that sign off on them (finance, underwriting, revenue cycle, operations). Without a named business owner for each definition, the layer drifts.

Should we build or buy a context layer?

Build if you have a strong data engineering team, a stable stack, and time; buy if speed matters and your systems are fragmented across vertical tools. Running pilots on top of unreconciled data is the most expensive option.

Sources

  • Anthropic Applied AI, "Effective context engineering for AI agents" (September 29, 2025): anthropic.com

  • Chroma Research, "Context Rot: How Increasing Input Tokens Impacts LLM Performance" (July 14, 2025): trychroma.com

  • Gartner, "Lack of AI-Ready Data Puts AI Projects at Risk" (February 26, 2025): gartner.com

  • Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027" (June 25, 2025): gartner.com

  • Computing, "MIT report: 95% of corporate generative AI pilots fail to deliver returns" (coverage of MIT NANDA, "The GenAI Divide: State of AI in Business 2025"): computing.co.uk

  • Anthropic, "Donating the Model Context Protocol and establishing the Agentic AI Foundation" (December 9, 2025): anthropic.com

  • LangChain, "State of Agent Engineering" (survey of 1,340 practitioners): langchain.com

  • The Decoder, "Shopify CEO and ex-OpenAI researcher agree that context engineering beats prompt engineering" (Karpathy definition): the-decoder.com

OutcomeCatalyst connects the systems you already run into a governed intelligence layer your team and your agents can reason over. Demos on this site use fictional data. To see this on your own operation, start a conversation.

Let’s build your company a brain
and put AI to work.

Let’s build your company a brain
and put AI to work.

Let’s build your company a brain
and put AI to work.

Your systems, your documents, and what your people have been carrying around in their heads. Thirty minutes to see what an AI brain could look like in your company.

Book a strategy call