HEALTHCARE · PATIENT DE-IDENTIFICATION

HIPAA expert determination for PHI and SUD data, built and attested by one team

Expert determination does not require a certified attester. It requires a qualified statistician and a method that holds up under scrutiny. OutcomeCatalyst builds the de-identification pipeline and writes the determination, so there is one team, one contract and one accountable party.

The problem

Why de-identified data still fails somebody else's review

These are the four things that stop a data set from being usable, and they tend to surface late, in a counterparty's legal review rather than in your own.

Your buyer's counsel cannot verify it
You removed the identifiers, but there is no written determination and no method log behind the work. The other side's lawyers have nothing to review, so the deal sits.
The attestation costs more than the pipeline
External attesters typically charge $75,000 to $100,000 for the determination report on its own, on top of whatever you already spent building the pipeline.
Standard de-identification does not cover Part 2
Records from federally assisted substance use disorder programs carry obligations HIPAA alone does not impose. Most providers will not take them on at all.
Safe Harbor removes the value you were selling
Stripping all eighteen identifier categories takes out the dates and geography research buyers need. The result is compliant and close to useless.
How this is usually bought

Two vendors, or one

The usual way to buy this splits the work between a pipeline builder and an attester. That split is where most of the cost and nearly all of the delay comes from.

The traditional model
A pipeline vendor builds the transformation, then hands it over
A separate attester assesses it and writes the determination
Two contracts, two scopes, two liability chains
Findings go back to the pipeline vendor for rework, then back again for re-review
Commonly six months or more from kickoff to a signed determination
The OutcomeCatalyst model
One team designs the pipeline against the risk model from the first day
The same team measures the risk and writes the determination
One contract, one scope, one accountable party
Rework happens inside the engagement instead of between two vendors
PHI and 42 CFR Part 2 handled together rather than farmed out separately
The pipeline and the determination are the same problem. Buying them from two vendors is what makes it slow.
Designing the transformation against the risk model is most of the work. Split it, and one party is guessing at what the other will accept.
Methodology

How the determination is actually reached

The standard is 45 CFR 164.514(b)(1): a qualified statistician applies generally accepted statistical and scientific principles and documents that the risk of re-identification is very small. Here is what that involves in practice.

1
Classify, then segment
Every structured field and every span of free text is classified as a direct identifier, a quasi-identifier, or neither. Records originating in federally assisted SUD programs are segmented at this point, before anything is transformed, because 42 CFR Part 2 governs them on a separate track.
2
Model the risk, not the field list
k-anonymity, l-diversity and t-closeness are measured across the quasi-identifiers that remain, and joint re-identification risk is modelled across the relational structure rather than field by field. Risk is evaluated against publicly available data, including census counts at whatever geographic level you want to keep.
3
Transform to a target, and keep what matters
Suppression and generalisation are applied where the measurements call for them and not elsewhere. This is why expert determination preserves more analytical value than Safe Harbor: dates and geography that research genuinely needs can survive when the numbers support keeping them.
4
Document, determine, then monitor
The method, the measurements and the residual risk are written up in full and the determination is signed. It runs for 24 months, with monitoring in between, because both the data set and the population around it keep moving.
What you receive

What the engagement produces

Five deliverables, from one team, under one contract.

The de-identification pipeline
Built in your environment or ours, and designed against the risk model rather than handed to someone else to assess afterwards.
The full expert determination report
Methodology, statistical analysis and the written determination, in the form a counterparty's counsel expects to be able to review.
A summary report you can share
A shorter document suitable for authorized recipients, so you are not handing the complete method log to every partner.
Twenty four months of validity
The determination holds for 24 months, with re-attestation available when the data set or the surrounding population changes materially.
Monthly monitoring
Ongoing checks between attestations, because a data set that keeps growing does not keep the same risk profile it had on the day it was signed.
Where it is used

What operators use this for

The same pipeline and the same determination support all of these.

Data commercialization
Releasing a data product to pharmaceutical and research buyers, with a determination their counsel can actually review before signing.
Internal analytics without PHI exposure
Giving analysts and data scientists a working data set without putting identified records in front of them.
Sharing under a data use agreement
Limited data sets under 164.514(e), where dates and geography are retained and the DUA carries the obligations.
Research collaborations
Multi-site and academic work where every party needs a defensible basis for what it received and what it may do with it.
Compliance workflows
Feeding SOC 2 and HITRUST readiness work with documented, repeatable de-identification instead of a one-off script nobody can explain.
Portfolio-wide data programs
Roll-ups and multi-site groups standardising de-identification across entities that each arrived with their own systems.
The systems it reads

Built for the stack a provider group actually runs

These are the systems referenced in the workflow above. Identifiers live in all of them, including the ones nobody thinks of as a source of PHI.

Epic
Encounters, notes and MRNs
Oracle Health
Inpatient records and orders
MEDITECH
Community hospital charts
athenahealth
Ambulatory charts and billing
Netsmart
SUD and behavioral health programs
Kipu and similar
Treatment program records
Clinical notes
Identifiers inline, in free text
Pathology and imaging
DICOM headers carry names
Lab results
HL7 feeds and ordering data
Claims 837 and 835
Members, subscribers and dates
Consent records
Part 2 consent and revocation
Datavant
Tokens for linkage after release
Try it on your own data

Try this on your own data before anyone else does

Take four fields you believe are safe, and try to find one patient.

Age, three-digit ZIP, diagnosis and visit month. Every one of them is permitted under Safe Harbor on its own. Put them together on an uncommon condition in a small population and you are often describing exactly one person.

Age
34
ZIP3
450
Diagnosis
1 in 2.1M
Visit month
March 2025

If you can single out a patient in your own data this way, that data set is not de-identified, whatever the field list says about it. Closing that gap is the entire reason expert determination exists.

Common questions

Questions operators and their counsel ask first

What is the difference between Safe Harbor and expert determination?

They are the two methods HIPAA permits. Safe Harbor, at 164.514(b)(2), removes eighteen categories of identifier and requires no actual knowledge that what remains could identify someone. Expert determination, at 164.514(b)(1), instead has a qualified statistician measure the actual risk and document that it is very small. Safe Harbor is simpler to apply. Expert determination usually leaves you with far more usable data.

Is there a certification for expert determination?

No. There is no certification, licence or official register of approved attesters, and anyone describing themselves as certified for this is describing something that does not exist. The rule asks for a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles for rendering information not individually identifiable, and for the methods and results to be documented. That is the whole bar, and it is a substantive one.

Who signs the determination?

The statistician who performed the assessment signs it, and their name and qualifications appear in the report. When you are evaluating any provider, including us, ask who will sign yours and what their background is. You should get a direct answer, because that person's judgment is what the determination rests on.

How long is a determination valid?

Ours runs for 24 months. HIPAA does not set a fixed expiry, but risk is measured against a data set and a population that both change, so an open-ended determination is not credible. We monitor monthly in between and re-attest when something moves enough to matter.

Can 42 CFR Part 2 SUD data be included?

Yes, and it is handled separately from the first step rather than folded in with everything else. Part 2 records are segmented before any transformation, and they are governed on their own track with their own consent and disclosure rules. In our default release design they are held out of external data products entirely and used only for consented or permitted research. This is the part most providers decline to take on.

Does de-identified data still count as PHI?

Once information meets the de-identification standard under 164.514, it is no longer PHI and the Privacy Rule no longer restricts its use or disclosure. That is exactly why the determination matters: it is the documented basis for saying the standard was met. A limited data set is a different thing and remains PHI, which is why it travels under a data use agreement.

How much does this cost?

It depends on the number of systems, how much of the identifying information sits in free text, and whether Part 2 records are in scope. The useful comparison is that external attesters typically charge $75,000 to $100,000 for the determination report alone, before anything is built. Because we do the pipeline and the determination together, the combined engagement generally comes in under that figure.

Can the data still be linked after release?

Yes, through privacy-preserving linkage tokens rather than by retaining identifiers. Records can be joined across sources and over time without anyone holding the underlying identity, which is what makes longitudinal research possible on a data set that is genuinely de-identified.

This page describes how we approach de-identification and expert determination under HIPAA and 42 CFR Part 2. It is general information about our methodology, not legal advice, and it does not replace your own counsel's review of a specific data set or disclosure.

Patient de-identification and expert determination: common questions

What is HIPAA expert determination?

It is one of the two de-identification methods HIPAA allows, set out at 45 CFR 164.514(b)(1). A qualified statistician applies accepted statistical and scientific principles, measures the risk that a record could be re-identified, and determines in writing that the risk is very small. The other method is Safe Harbor, which removes eighteen categories of identifier.

Do you need to be certified to provide an expert determination?

No. There is no certification, licence or register of approved attesters for this. The rule asks for a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles, and for the methods and results to be documented. Operators frequently pay more than they need to because they assume a certification is involved.

Can 42 CFR Part 2 SUD data be de-identified?

Yes, and it is handled on its own track from the start. Records from federally assisted substance use disorder programs carry obligations that HIPAA alone does not impose, so they are segmented before any transformation and governed separately. Many providers of de-identification will not take on Part 2 data at all.

Unified operating layer to harness artificial intelligence. Connect fragmented data, create agentic workflows, enable faster decisions across your company.

© 2026 OutcomeCatalyst. All rights reserved.

Unified operating layer to harness artificial intelligence. Connect fragmented data, create agentic workflows, enable faster decisions across your company.

© 2026 OutcomeCatalyst. All rights reserved.

Unified operating layer to harness artificial intelligence. Connect fragmented data, create agentic workflows, enable faster decisions across your company.

© 2026 OutcomeCatalyst. All rights reserved.