‹ Back to Blog

Commercial Real Estate

Why Your Yardi, MRI, and ARGUS Numbers Never Match: The CRE Data Layer Explained (2026)

Zach Shapiro

·

·

14 min read

Modern glass commercial office building representing a commercial real estate portfolio asset

TL;DR: Yardi, MRI, and ARGUS disagree because each was built for a different job, holds a different slice of the truth, updates on its own clock, and shares no identifier for a property, suite, or tenant. Occupancy alone has four defensible definitions. A CRE data layer resolves properties and tenants into persistent identities, holds one canonical set of metric definitions, and reads the lease and loan documents where the terms that drive cash flow actually live. This guide covers why the numbers diverge, how entity resolution works on CRE data, what AI can and cannot extract from a rent roll, and a realistic 90-day sequence.

If you have ever sat in an asset management meeting where the property management system says occupancy is 91.4%, the underwriting model says 89%, and the investor report says 92%, you already understand the problem this article is about. All three numbers are defensible. They are computed from different sources, on different dates, using different definitions of what counts as occupied. Nothing is broken. Nothing is reconciled either.

Commercial real estate has an unusually severe version of this problem because the industry's core systems were built for different jobs in different decades and were never designed to agree with each other. Argus Enterprise underwrites. Yardi and MRI operate and account. VTS tracks leasing activity. CoStar prices the market. Thousands of lease PDFs hold the terms that determine what any of those numbers actually mean. The same property is spelled three ways, the same tenant appears under a DBA in one system and a legal entity in another, and the suite numbering does not match between the rent roll and the lease file.

This guide explains why that happens, what a CRE data layer is, how entity resolution works on properties and tenants, what AI can and cannot extract from a rent roll, and a realistic sequence for fixing it. It is written for owners, principals, heads of acquisitions, and asset management leaders.

Key takeaways

  • Deloitte's 2026 Commercial Real Estate Outlook surveyed more than 850 C-level executives and direct reports at firms with at least $250 million in AUM across 13 countries. 19% said they remain in the early stages of their AI journey, and 27% reported challenges with AI implementation including technical issues, lack of expertise, and internal resistance.

  • Only 22% leverage industry-specific software platforms such as integrated workplace management systems, and 20% use publicly available large language models. The gap between those two numbers is where most CRE AI work currently sits: general-purpose models pointed at industry-specific data they cannot resolve.

  • Roughly half of Deloitte's respondents flagged generating synthetic data as an area of high interest, which is a revealing signal. Interest in synthetic data is usually a symptom of real data that is too fragmented or too incomplete to use.

  • Vendors now claim sub-15-second extraction of rent rolls and T12s at around 99% field-level accuracy. Extraction is largely solved. Reconciliation across systems is not, and that is the harder half.

  • The knowledge graph market is projected to grow from about $1.9 billion in 2026 to roughly $9.88 billion by 2032, a 31.6% CAGR, per MarketsandMarkets. The driver is that agents cannot operate reliably without governed context about how a business actually works.

  • Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027. In CRE the most common reason is that the agent was deployed on top of systems that do not agree.

Why do Yardi, MRI, and ARGUS produce different numbers for the same property?

Because each system was designed to answer a different question, holds a different subset of the truth, and updates on a different clock. There is no shared identifier for a property, a tenant, or a lease across them, so nothing forces agreement.

Here is where the divergence actually comes from.

Different definitions of the same metric. Occupancy can be measured on physical occupancy, leased occupancy, economic occupancy, or occupancy net of tenants in abatement. A property that is 100% leased with two tenants in free rent is somewhere between 87% and 100% occupied depending on which definition you use. Both the property manager and the asset manager are right.

Different measurement standards for area. Gross leasable area, net rentable area, and usable area differ, and load factors vary by building and sometimes by lease. Rent per square foot is not comparable across a portfolio until you know which denominator each system used.

Different treatment of recoveries. CAM, taxes, and insurance recoveries are estimated monthly and reconciled annually. The accounting system holds the estimate during the year, the reconciliation adjusts it after year end, and the underwriting model holds an assumption that was set at acquisition. Three different numbers for the same line, all correct at their own moment.

Different timing. Yardi or MRI reflects posted transactions as of the last close. VTS reflects leasing activity as of this morning. Argus reflects the underwriting as of the last model update, which may be eleven months old. CoStar reflects market data on its own refresh cycle.

No shared entity identity. This is the deepest cause. "Ridgeline Logistics Center," "Ridgeline Log. Ctr.," and "RLC Building A" are three strings in three systems and one asset in reality. "Acme Distribution LLC" signs the lease, "Acme Dist." pays the rent, and "Acme Group" is the guarantor. Nothing in the software stack knows those are related.

Lease terms trapped in prose. The percentage rent breakpoint, the co-tenancy clause, the expansion option, the termination right, the escalation basis: these determine the cash flow and they live in a PDF. A model that does not read the lease is modeling an assumption about the lease.

Any one of these produces a variance. All six together produce a portfolio where nobody can say what NOI is without a two-week reconciliation exercise.

What is a CRE data layer?

A CRE data layer is a governed layer between your operating systems and your decisions that does three things: connects to Yardi, MRI, Argus, VTS, CoStar, your accounting system, and your document repositories; resolves the entities those systems describe (properties, units, tenants, leases, legal entities, counterparties) into single persistent identities; and holds the canonical definition of every metric your firm reports on, so occupancy means one thing across the portfolio.

It is not a replacement for your systems. It is the thing that has been missing between them. Yardi keeps operating the properties. Argus keeps underwriting. The layer is what lets you ask a question that spans both and get one answer with lineage back to the source records.

The reason to build it as a graph rather than a set of joined tables is that CRE relationships are the analysis. A tenant belongs to a parent guarantor, occupies three suites across two properties in one MSA, has a lease with an expansion option that encumbers a suite currently occupied by someone else, and appears in a fund that has a preferred return owed to an LP. Those are edges, and edges are what you actually underwrite. We described the general architecture in Knowledge Graphs for Enterprise AI and the buyer-facing version in What Is an AI Context Layer?.

What is entity resolution in commercial real estate, and why is property and tenant matching so hard?

Entity resolution is the process of determining that different records across different systems refer to the same real-world property, tenant, or lease, and then maintaining one persistent identity for it even as names, ownership, and attributes change.

The standard pipeline ingests records continuously, standardizes them against defined rules, matches on exact identifiers where they exist and fuzzy logic where they do not, merges survivors into a golden record with lineage, and persists that identity through change.

CRE makes each step harder than average for specific reasons.

Properties have no universal key. There is a parcel or APN, a street address, an internal property code per system, a CoStar ID, and a name that marketing changes. Addresses are inconsistently formatted, and multi-parcel assets and campuses break one-to-one assumptions. A single asset can legitimately have four parcels and two addresses.

Suites and units get renumbered. A demised suite becomes 210A and 210B. A combined suite becomes 300. The rent roll reflects the new numbering, the lease file reflects the old, and the historical performance series silently breaks.

Tenants operate under layered identities. The legal entity on the lease, the DBA on the signage, the parent guarantor on the credit, the payment entity on the remittance. Credit exposure is a function of the parent, occupancy is a function of the suite, and collections are a function of the payment entity. Roll them up wrong and your concentration analysis is wrong.

Ownership structures nest. Property to SPE to JV to fund to sponsor, with promote structures and partner splits at multiple levels. Attributing NOI to an ownership share requires the ownership graph, not a spreadsheet column.

Time matters. Entity resolution in CRE has to be temporal. Who owned this asset in Q3 2024, which tenant occupied suite 210 before the demise, what was the escalation basis before the amendment. A model that only knows current state cannot explain a variance.

Get this right and the questions that were previously projects become queries. Our NOI and occupancy intelligence demo shows occupancy and lease expiry across a portfolio on a resolved model, using fictional data.

Can AI actually read a rent roll, T12, or offering memorandum accurately?

Yes, and this is the part of the problem that has genuinely been solved. Purpose-built CRE extraction tools now pull structured data from rent rolls, trailing twelve-month operating statements, and offering memoranda in seconds, with vendors reporting field-level accuracy around 99% on standard formats regardless of broker layout. Document-specific vision models handle the table and chart parsing that used to be the bottleneck.

Three caveats matter more than the accuracy number.

Accuracy on standard formats is not accuracy on your worst broker's spreadsheet. The 99% figures are vendor-reported on typical documents. A rent roll exported from a regional owner's twelve-year-old system with merged cells and footnoted abatements performs worse. Test on your ugliest five files, not on a clean sample.

Extraction produces fields, not judgment. The extractor reads "$28.50/SF NNN, 3% annual escalation, 2 months free." It does not decide whether the free rent belongs in year one effective rent under your firm's convention, whether the escalation compounds or is simple, or whether the reported operating expenses are grossed up to a normalized occupancy. Those are your definitions, and they belong in the layer.

Extraction without resolution creates a new silo. If extracted lease data lands in its own tool with its own tenant list, you have added a fourth system that disagrees with the other three. The extraction has to write into a resolved model or it makes the reconciliation problem worse.

The useful sequence is extract, then resolve, then apply canonical definitions, then underwrite. Most firms buy the first step and skip the middle two. Our underwriting acceleration demo walks through a broker pro forma compared against reality on the same asset, and the deal origination demo shows parcel-level signal scoring. Both use fictional data, and the full set of workflows sits on our commercial real estate page.

Do I need to replace Yardi, MRI, or ARGUS?

No. These platforms remain the industry standard for property operations, accounting, and institutional underwriting, and replacing them is a multi-year project with real operational risk. They are also, by design, not integration layers.

What is fair to say is that they were built before modern AI, and adding AI features to a legacy codebase is a different exercise from building on a connected data model. Argus still requires lease-by-lease data entry for a reason: its model structure predates the idea that lease data would arrive from anywhere else.

The practical stance is to leave the systems of record in place and add the layer that makes them agree. That means Yardi stays authoritative for posted financials, Argus stays authoritative for the underwriting model, VTS stays authoritative for pipeline, and the layer holds the resolved entities and canonical definitions that let you compare them. When someone asks why occupancy differs between two reports, the answer becomes a two-minute lineage trace instead of a two-week exercise.

The build-versus-buy economics on the layer itself are covered in Buy vs. Build: The AI Context Layer Decision.

How do you reconcile projected NOI against actual NOI?

You reconcile it by decomposing the variance into named, attributable drivers on a resolved model, rather than comparing two totals and arguing. A projected-versus-actual comparison at the NOI line tells you the size of the problem and nothing about its cause.

A working decomposition separates:

  1. Rent variance from occupancy (units vacant that were modeled occupied, or the reverse), measured at the suite level with dates.

  2. Rent variance from rate (contract rent achieved versus modeled, including free rent timing and escalation basis differences).

  3. Recovery variance (estimated recoveries versus the annual reconciliation, plus gross-up assumptions and any caps or base-year mechanics that the model treated differently from the lease).

  4. Other income variance (parking, percentage rent against actual sales reporting, storage, fees).

  5. Operating expense variance by category, separating controllable from non-controllable, with tax reassessments and insurance renewals isolated because they are structurally different from managed expenses.

  6. Timing and one-time items (capital treated as expense, prior-period adjustments, tenant reimbursement true-ups).

Every line traces to source records: a lease clause, a posted transaction, a rent roll row, a tax bill. Once this decomposition runs automatically each month rather than being rebuilt by an analyst each quarter, asset management shifts from explaining variances to preventing them.

What systems does a CRE data layer need to connect?

Six categories, and the last two are the ones firms skip.

Property management and accounting. Yardi, MRI, RealPage, AppFolio, Entrata. Source of posted financials, rent roll, delinquency, and recovery billing.

Underwriting and valuation. Argus Enterprise, Excel models, appraisal files. Source of assumptions, hold-period cash flows, and the exit math.

Leasing and pipeline. VTS, Salesforce or HubSpot for the deal side, broker communications. Source of pipeline, tour activity, and pending deals not yet in the accounting system.

Market and property data. CoStar, Crexi, Reonomy, county records, tax assessor data. Source of comps, ownership, parcel detail, and market rent signals.

Documents. Leases and amendments, estoppels, SNDAs, loan agreements, JV agreements, PSAs, service contracts, tax bills, insurance policies. This is where the terms that drive cash flow actually live.

Loan and capital stack systems. Debt schedules, covenant terms, reserve accounts, lender reporting requirements. A DSCR covenant projection that does not read the loan agreement is an estimate.

The document and loan categories are where the highest-value answers hide, and they are the two that no dashboard vendor connects for you.

What can AI agents do once the layer exists?

They can do the reading, comparing, and drafting that currently consumes analyst weeks. Concretely:

  • First-pass underwriting from an OM. Extract, resolve against your market data, apply your firm's underwriting conventions rather than the broker's, and flag every assumption where the OM is more optimistic than your comps support.

  • Lease abstraction with obligation tracking. Not just extracting terms but placing them on a calendar: expirations, options, notice deadlines, escalation dates, co-tenancy triggers. Missed notice deadlines are a recurring, entirely preventable loss.

  • Monthly NOI variance narrative. The decomposition above, drafted with sources attached, ready for asset management review.

  • Portfolio exposure queries. Lease expiry concentration by MSA and quarter, top parent-level tenant exposure across funds, rollover risk against market rent by submarket.

  • Deal triage at volume. Screening dozens of OMs against thesis criteria so analysts underwrite the five that merit it. See AI Agents, Workflows, and Knowledge Graphs in Commercial Real Estate for the broader landscape.

What they should not do unsupervised: set the mark, sign off on a covenant certification, or produce an investor-facing number without lineage. The value comes from removing the reading and reconciling labor, not from removing the judgment.

What does a realistic 90-day sequence look like?

Days 1 to 15: define and scope. Write down your firm's canonical definitions for the twenty metrics that appear in your investor reporting and IC memos: occupancy, effective rent, NOI, recoveries, and the rest. Pick one asset class and five to ten assets. Mixing office, industrial, and multifamily in a pilot triples the definitional work.

Days 16 to 45: connect and extract. Connect the property management and accounting system, the underwriting models, and the leasing pipeline for those assets. Load the full lease file including amendments, plus the loan documents. Run extraction and review the output field by field on the five worst documents.

Days 46 to 70: resolve and reconcile. Run entity resolution across properties, suites, tenants, and parent entities. Then reconcile the layer's output against the property accountant's last close, line by line, until they agree or the difference is explained and documented. This step is what makes the number authoritative internally.

Days 71 to 90: put one decision on it. The monthly asset management review, the quarterly investor reporting pack, or the acquisitions screening queue. Measure the before and after cycle time.

Then expand by asset class rather than by asset count, because each new asset class introduces new definitional work while each new asset in a known class is nearly free.

What goes wrong

Buying extraction and calling it a data layer. Extraction is one step of five. Extracted data that lands in its own silo adds a system to reconcile.

Letting each fund or region keep its own definitions. If the East portfolio measures economic occupancy and the West measures leased occupancy, portfolio-level occupancy is meaningless. Someone has to decide, and it has to be someone senior enough that the decision holds.

Skipping the lease file. Firms connect the accounting system because it has an API and defer the leases because they are PDFs. The leases hold the terms that explain the variances, so this ordering guarantees the layer cannot answer the questions that motivated it.

Ignoring temporal state. Building a current-state-only model makes historical performance analysis impossible after the first suite renumbering or ownership change.

No named owner. Definitions drift, feeds break, and mappings go stale after a system upgrade. Someone owns it or it decays within two quarters.

How do you measure whether it worked?

  • Days from month-end close to a reconciled portfolio NOI view. Many firms baseline at three to six weeks.

  • Analyst hours per underwriting. Track the full path from OM received to IC-ready model.

  • Deal throughput. Number of opportunities screened per month at constant headcount. Origination advantage in a competitive market is largely a function of how many deals you can look at seriously.

  • Missed deadline count. Lease options, notice dates, and covenant certifications missed per year. Target zero, and measure it.

  • Variance explanation time. How long it takes to answer "why is NOI off by $180K" with sources.

Common questions about CRE data reconciliation

Why don't Yardi and ARGUS match?

They answer different questions on different clocks. Yardi holds posted transactions as of the last accounting close; Argus holds underwriting assumptions as of the last model update. They also lack a shared identifier for properties, suites, and tenants, so nothing forces them to agree. Reconciliation requires resolving entities between them and applying one canonical set of metric definitions.

What is a CRE data layer?

A governed layer that connects your property management, accounting, underwriting, leasing, market data, and document systems, resolves properties and tenants into single persistent identities, and holds your firm's canonical metric definitions so every report computes the same way.

Can AI extract data from a rent roll or T12?

Yes. Purpose-built CRE extraction tools process rent rolls, T12s, and offering memoranda in seconds with vendor-reported accuracy near 99% on standard formats. The remaining work is resolving that extracted data against your existing portfolio and applying your own underwriting conventions.

Do I need a data warehouse for commercial real estate?

A warehouse is a useful storage substrate but is not sufficient by itself. It stores rows without knowing that two property records are the same asset, which occupancy definition is authoritative, or what a lease clause means. Those functions sit in the layer above the warehouse.

How do you handle tenants that operate under multiple entities?

Model the legal entity, the DBA, the payment entity, and the parent guarantor as distinct nodes with explicit relationships, then roll up differently depending on the question: credit exposure to the parent, occupancy to the suite-level entity, collections to the payment entity.

What about temporal changes like suite demises and ownership transfers?

The model has to be time-aware, storing which entity occupied which space during which period and who owned the asset when. Current-state-only models break historical comparisons the first time a suite is renumbered.

How long does this take?

A five to ten asset pilot in a single asset class is realistic in about 90 days, including reconciliation with your property accountant. Expansion is faster within the same asset class and slower across new ones.

Does this replace my asset management team's judgment?

No. It removes the reconciliation and reading labor that currently consumes most of their week, which is what makes room for judgment. The firms getting value are redeploying analyst time, not reducing headcount.

How does this apply to a smaller portfolio?

The definitional work scales down cleanly; the entity resolution work does not disappear but is smaller. Owners with fewer than ten assets often get most of the value from lease abstraction with obligation tracking plus a single canonical NOI definition, and can defer the rest.

Sources

  • Deloitte, "2026 Commercial Real Estate Outlook" (survey of more than 850 C-level executives and direct reports at firms with $250M+ AUM across 13 countries, fielded June to July 2025; 19% early stage AI, 27% implementation challenges, 22% industry-specific platforms, 20% public LLMs, synthetic data interest): deloitte.com

  • MarketsandMarkets, "Knowledge Graph Market Report" (approximately $1.9B in 2026 to $9.88B by 2032, 31.6% CAGR): marketsandmarkets.com

  • Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027" (June 2025): gartner.com

  • Docsumo, CRE underwriting document extraction (rent rolls, T12s, and operating statements), vendor-reported: docsumo.com

  • RealQuant, CRE document intelligence for offering memoranda, rent rolls, and T-12s, vendor-reported accuracy on standard formats: realquant.ai

  • Senzing, "What Is Entity Resolution? How It Works and Why It Matters" (ingest, standardize, match, merge, persist pipeline; golden record concept): senzing.com

OutcomeCatalyst connects the systems you already run into a governed intelligence layer your team and your agents can reason over. Demos on this site use fictional data. To see this on your own portfolio, start a conversation.

Unified operating layer to harness artificial intelligence. Connect fragmented data, create agentic workflows, enable faster decisions across your company.

© 2026 OutcomeCatalyst. All rights reserved.