HEALTHCARE · PATIENT DE-IDENTIFICATION

HIPAA expert determination for PHI and SUD data, built and attested by one team

Expert determination does not require a certified attester. It requires a qualified statistician and a method that holds up under scrutiny. OutcomeCatalyst builds the de-identification pipeline and writes the determination, so there is one team, one contract and one accountable party.

The problem

Why de-identified data still fails somebody else's review

These are the four things that stop a data set from being usable, and they tend to surface late, in a counterparty's legal review rather than in your own.

Your buyer's counsel cannot verify it
You removed the identifiers, but there is no written determination and no method log behind the work. The other side's lawyers have nothing to review, so the deal sits.
The determination is a snapshot, and the data keeps moving
A determination describes the data set as it stood on the day it was signed. Records keep arriving and the population around them shifts, and nothing tells you when the basis you relied on has stopped holding.
Standard de-identification does not cover Part 2
Records from federally assisted substance use disorder programs carry obligations HIPAA alone does not impose. Most providers will not take them on at all.
Safe Harbor removes the value you were selling
Stripping all eighteen identifier categories takes out the dates and geography research buyers need. The result is compliant and close to useless.
How this is usually bought

Two vendors, or one

The usual way to buy this splits the work between a pipeline builder and an attester. That split is where the delay comes from, and it is why neither party ends up owning the result.

The traditional model
A pipeline vendor builds the transformation, then hands it over
A separate attester assesses it and writes the determination
Two contracts, two scopes, two liability chains
Findings go back to the pipeline vendor for rework, then back again for re-review
Commonly six months or more from kickoff to a signed determination
The OutcomeCatalyst model
One team designs the pipeline against the risk model from the first day
The same team measures the risk and writes the determination
One contract, one scope, one accountable party
Rework happens inside the engagement instead of between two vendors
PHI and 42 CFR Part 2 handled together rather than farmed out separately
The pipeline and the determination are the same problem. Buying them from two vendors is what makes it slow.
Designing the transformation against the risk model is most of the work. Split it, and one party is guessing at what the other will accept.
Methodology

How the determination is actually reached

The standard is 45 CFR 164.514(b)(1): a qualified statistician applies generally accepted statistical and scientific principles and documents that the risk of re-identification is very small. Here is what that involves in practice.

1
Classify, then segment
Every structured field and every span of free text is classified as a direct identifier, a quasi-identifier, or neither. Records originating in federally assisted SUD programs are segmented at this point, before anything is transformed, because 42 CFR Part 2 governs them on a separate track.
2
Measure the risk against numbers, not a checklist
Three measurements, each with a threshold the release has to clear before it moves. No record may match fewer than five people in the wider population, and no fewer than ten where the record came from a Part 2 program. Any group that size has to hold at least three different diagnoses, so finding the group does not hand over the condition. And the mix of values inside a group has to stay close to the mix across the whole release. These are k-anonymity, l-diversity and t-closeness, and they are measured across the full linked patient record at production volume rather than one table at a time, because that is how somebody trying to re-identify a person would actually work.
3
Change only what the numbers say to change
Suppression and generalisation are applied where a measurement misses its threshold and nowhere else, which is why this preserves so much more than Safe Harbor. Dates are shifted per patient by an offset that holds the gaps between visits intact, so length of stay and time to follow-up survive even though the calendar dates do not.
4
The evidence is produced, then a person signs it
The pipeline produces the measurements. A qualified statistician reads them and signs the determination under 164.514(b)(1), and nothing leaves as de-identified until that signature exists. If a threshold is missed the release stops rather than going out with a caveat attached. The signed determination is tied to the exact data set by a checksum, so if one byte changes it no longer applies and the data has to be measured again.
What you receive

What the engagement produces

Five deliverables, from one team, under one contract.

The de-identification pipeline
Built in your environment or ours, and designed against the risk model rather than handed to someone else to assess afterwards.
The full expert determination report
Methodology, statistical analysis and the written determination, in the form a counterparty's counsel expects to review, with the checksum that binds it to the exact data set it describes.
A summary report you can share
A shorter document suitable for authorized recipients, so you are not handing the complete method log to every partner.
Twenty four months of validity
The determination holds for 24 months, with re-attestation available when the data set or the surrounding population changes materially.
Monthly monitoring
Ongoing checks between attestations, because a data set that keeps growing does not keep the same risk profile it had on the day it was signed.
Where it is used

What operators use this for

One pipeline, but not one legal basis. The expert determination covers the de-identified releases. A limited data set is a different thing and travels under its own agreement, so each use below gets the release it actually needs.

De-identified, under the determination
Data commercialization
Releasing a data product to pharmaceutical and research buyers, with a determination their counsel can actually review before signing.
De-identified, under the determination
Internal analytics without PHI exposure
Giving analysts and data scientists a working data set without putting identified records in front of them.
Limited data set, still PHI
Sharing under a data use agreement
A limited data set keeps dates and geography, so it stays PHI and the determination does not cover it. It moves under a data use agreement at 164.514(e), which carries its own obligations for whoever receives it.
Depends what the study needs
Research collaborations
De-identified where the protocol allows it, a limited data set under a DUA where the study genuinely needs dates and geography. The difference is decided before anything moves.
De-identified, under the determination
Compliance workflows
Feeding SOC 2 and HITRUST readiness work with documented, repeatable de-identification instead of a one-off script nobody can explain.
De-identified, under the determination
Portfolio-wide data programs
Roll-ups and multi-site groups standardising de-identification across entities that each arrived with their own systems.
The systems it reads

Built for the stack a provider group actually runs

These are the systems referenced in the workflow above. Identifiers live in all of them, including the ones nobody thinks of as a source of PHI.

Epic
Encounters, notes and MRNs
Oracle Health
Inpatient records and orders
MEDITECH
Community hospital charts
athenahealth
Ambulatory charts and billing
Netsmart
SUD and behavioral health programs
Kipu and similar
Treatment program records
Clinical notes
Identifiers inline, in free text
Pathology and imaging
DICOM headers carry names
Lab results
HL7 feeds and ordering data
Claims 837 and 835
Members, subscribers and dates
Consent records
Part 2 consent and revocation
Datavant
Tokens for linkage after release
Try it on your own data

Try this on your own data before anyone else does

Take four fields and try to find one patient.

Age, three-digit ZIP and diagnosis are all permitted under Safe Harbor. The visit month is not: Safe Harbor removes every date element below the year. Expert determination is the only method that can keep finer timing at all, and only for the records where the numbers show it is safe to. Put these four together on an uncommon condition in a small population and you are often describing exactly one person.

Age
34
Permitted
ZIP3
450
Permitted
Diagnosis
1 in 2.1M
Permitted
Visit month
March 2025
Safe Harbor cuts this to the year

If you can single out a patient in your own data this way, that data set is not de-identified, whatever the field list says about it. A checklist cannot tell you this, because it never looks at your data. Measuring it is the whole reason expert determination exists.

Common questions

Questions operators and their counsel ask first

What is the difference between Safe Harbor and expert determination?

They are the two methods HIPAA permits. Safe Harbor, at 164.514(b)(2), removes eighteen categories of identifier and requires no actual knowledge that what remains could identify someone. Expert determination, at 164.514(b)(1), instead has a qualified statistician measure the actual risk and document that it is very small. Safe Harbor is simpler to apply. Expert determination usually leaves you with far more usable data.

Is there a certification for expert determination?

No. There is no certification, licence or official register of approved attesters, and anyone describing themselves as certified for this is describing something that does not exist. The rule asks for a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles for rendering information not individually identifiable, and for the methods and results to be documented. That is the whole bar, and it is a substantive one.

Who signs the determination?

The statistician who performed the assessment signs it, and their name and qualifications appear in the report. When you are evaluating any provider, including us, ask who will sign yours and what their background is. You should get a direct answer, because that person's judgment is what the determination rests on.

Can we keep dates and geography?

Sometimes, and this is where expert determination earns its keep. Safe Harbor removes every date element below the year and all geography below the state, which is exactly what makes it so expensive for research. Expert determination can justify keeping finer detail, but only where the measurements show the risk stays very small, and only for the records where that holds. Everywhere else the detail is generalised or the record comes out. If you need dates and geography kept across the board no matter what, that is a limited data set under a data use agreement, not de-identified data.

What exactly gets measured?

Three things, each against a threshold. That no record in the release matches fewer than five people in the wider population, and no fewer than ten for records from Part 2 programs. That any group that size holds at least three different diagnoses. And that the mix of values inside a group stays close to the mix across the release as a whole. The measurements run across the full linked patient record rather than table by table, because a row that looks safe on its own often is not once it is joined back to everything else about that patient.

How long is a determination valid?

Ours runs for 24 months. HIPAA does not set a fixed expiry, but risk is measured against a data set and a population that both change, so an open-ended determination is not credible. We monitor monthly in between and re-attest when something moves enough to matter.

Can 42 CFR Part 2 SUD data be included?

Yes, and it is handled separately from the first step rather than folded in with everything else. Part 2 records are segmented before any transformation, and they are governed on their own track with their own consent and disclosure rules. In our default release design they are held out of external data products entirely and used only for consented or permitted research. This is the part most providers decline to take on.

Does de-identified data still count as PHI?

Once information meets the de-identification standard at 164.514, it is no longer PHI and the Privacy Rule no longer restricts its use or disclosure. That is precisely why the determination matters: it is the documented basis for saying the standard was met. A limited data set is a different thing. It keeps dates and geography, it remains PHI, the determination does not cover it, and it travels under a data use agreement instead.

How much does this cost?

It depends on how many systems are in scope, how much of the identifying information sits in free text rather than in structured fields, and whether Part 2 records are included. Because one team builds the pipeline and signs the determination, you are buying a single engagement rather than two, with no coordination overhead between vendors and no rework handed back and forth. We scope it against your actual data rather than quoting from a rate card, so the first step is a look at what you are holding.

Can the data still be linked after release?

Yes, through privacy-preserving linkage tokens rather than by retaining identifiers. Records can be joined across sources and over time without anyone holding the underlying identity, which is what makes longitudinal research possible on a data set that is genuinely de-identified.

This page describes how we approach de-identification and expert determination under HIPAA and 42 CFR Part 2. It is general information about our methodology, not legal advice, and it does not replace your own counsel's review of a specific data set or disclosure.

Patient de-identification and expert determination: common questions

What is HIPAA expert determination?

It is one of the two de-identification methods HIPAA allows, set out at 45 CFR 164.514(b)(1). A qualified statistician applies accepted statistical and scientific principles, measures the risk that a record could be re-identified, and determines in writing that the risk is very small. The other method is Safe Harbor, which removes eighteen categories of identifier.

Do you need to be certified to provide an expert determination?

No. There is no certification, licence or register of approved attesters for this. The rule asks for a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles, and for the methods and results to be documented. Operators frequently pay more than they need to because they assume a certification is involved.

Can 42 CFR Part 2 SUD data be de-identified?

Yes, and it is handled on its own track from the start. Records from federally assisted substance use disorder programs carry obligations that HIPAA alone does not impose, so they are segmented before any transformation and governed separately. Many providers of de-identification will not take on Part 2 data at all.

Unified operating layer to harness artificial intelligence. Connect fragmented data, create agentic workflows, enable faster decisions across your company.

© 2026 OutcomeCatalyst. All rights reserved.

Unified operating layer to harness artificial intelligence. Connect fragmented data, create agentic workflows, enable faster decisions across your company.

© 2026 OutcomeCatalyst. All rights reserved.

Unified operating layer to harness artificial intelligence. Connect fragmented data, create agentic workflows, enable faster decisions across your company.

© 2026 OutcomeCatalyst. All rights reserved.