← Back to Blog
Engineering·11 min read·Sep 28, 2026

De-Identifying Dental Records for Model Evaluation: HIPAA Safe Harbor vs Expert Determination in Practice

De-Identifying Dental Records for Model Evaluation: HIPAA Safe Harbor vs Expert Determination in Practice

In 2000, Latanya Sweeney's research at Carnegie Mellon estimated that approximately 87% of the U.S. population could be uniquely identified by just three fields: five-digit ZIP code, gender, and date of birth. All three fields sit in a standard dental practice-management export, usually within the first ten columns of the patient table.

That said, a clinical AI team can't validate a charting model, a note summarizer, or a radiograph classifier against toy synthetic data and expect the numbers to hold in production. The eval set has to look like real charts, real notes, and real images — so someone has to strip PHI out of real records before they reach an eval pipeline.

This post covers the two de-identification paths HIPAA recognizes and what each one removes from dental charts, clinical notes, and radiograph metadata. It also covers the places where free-text notes keep leaking identifiers after the structured fields come back clean.

Why Eval Pipelines Are A PHI Problem First

Eval data spreads further than training data does. A single eval run can copy records into an S3 bucket, send them to a model endpoint, write them into a tracing tool, and paste failure cases into a Slack thread so an engineer can triage them.

Each of those surfaces needs to sit inside your BAA boundary if the records are PHI. In practice, several of them do not, which is why we treat de-identification as a precondition for any eval work rather than a cleanup step. Our overview of HIPAA requirements for clinical AI in dental practices covers the boundary itself in more detail.

HIPAA offers two de-identification methods. Safe Harbor removes 18 listed identifiers. Expert Determination has a qualified expert certify that re-identification risk is very small. Once either method is met, the data is no longer PHI.

Keep in mind that de-identified data falls outside the Privacy Rule entirely. This is why the method matters so much: it decides whether your eval harness, your trace store, and your vendor's logging layer need to be HIPAA-eligible at all.

What Safe Harbor Actually Requires

Safe Harbor, defined at 45 CFR 164.514(b)(2), is the prescriptive path. You remove 18 categories of identifiers belonging to the patient or to their relatives, employers, or household members. The covered entity must also have no actual knowledge that the remaining information could identify the individual.

The 18 identifiers group into six practical buckets for a dental dataset. Each bucket includes, but is not limited to, the following:

  • Names and contact details. Names, telephone numbers, fax numbers, email addresses, URLs, and IP addresses. Patient portal logs and online booking tables carry the last two more often than teams expect.
  • Geography. Every subdivision smaller than a state, including street address, city, county, and ZIP code. The first three ZIP digits may stay only if that combined area holds more than 20,000 people; otherwise they become 000.
  • Dates and ages. Every element of a date except the year for dates tied to the individual, including birth, visit, and procedure dates. Ages over 89 collapse into a single category of 90 or older.
  • Account and record numbers. Social Security numbers, medical record numbers, health plan beneficiary numbers, account numbers, and certificate or license numbers. In dental data, the insurance subscriber ID is the one most often missed.
  • Devices and vehicles. Vehicle identifiers including license plates, plus device identifiers and serial numbers. Implant lot numbers and sensor serial numbers both fall here.
  • Biometrics, images, and catch-all codes. Finger and voice prints, full-face photographs and any comparable images, and any other unique identifying number, characteristic, or code. That last category is the one auditors press on.

Under HHS guidance based on 2000 Census data, 17 three-digit ZIP prefixes fall below the 20,000-person threshold and must be zeroed. Remember that the list depends on the census, so your geography step needs to reference the current HHS guidance rather than a hardcoded array someone wrote in 2019.

The actual-knowledge clause. Safe Harbor still fails if your team knows the remaining data could identify someone. For example, if a note says the patient is "the only orthodontist in town," stripping the 18 fields does not cure it.

What Expert Determination Actually Requires

Expert Determination, at 45 CFR 164.514(b)(1), replaces the checklist with a judgment. A person with appropriate knowledge of statistical and scientific methods determines that the risk is very small that an anticipated recipient could identify an individual, using the data alone or combined with other reasonably available information.

That judgment rests on specific commitments that the expert documents and the covered entity keeps on file. They include:

  • A named recipient and environment. Risk is assessed against who receives the data and what controls surround it. An internal eval cluster behind IAM and VPC controls scores very differently from a public benchmark release.
  • Quantified risk. The expert measures quasi-identifier combinations, such as age band, three-digit ZIP, sex, and rare procedure codes, against equivalence-class thresholds. They then generalize or suppress fields until the risk clears.
  • Written methods and results. HIPAA requires the expert to document the analysis and its justification. That document is what you produce during an OCR inquiry or a customer security review.
  • Scope limits. Experts commonly bind the determination to a dataset version and a time window. A new data pull or a new recipient calls for re-review.

The trade is cost for fidelity. Engagements are typically priced in the five figures per dataset, depending on complexity and on whether free text is in scope. In return, you get to keep information that Safe Harbor would destroy.

DimensionSafe HarborExpert Determination
Legal basis45 CFR 164.514(b)(2)45 CFR 164.514(b)(1)
Visit and procedure datesYear onlyPer-patient date shift usually permitted
GeographyState, or three-digit ZIP above 20,000 peopleThree-digit ZIP or region if the risk analysis supports it
Ages90+ collapsedAge bands chosen by the expert
Free-text notesAll 18 identifiers removed from the textScrubbed, with measured residual risk
DocumentationInternal procedure recordFormal written determination
Typical costEngineering timeEngineering time plus a five-figure expert engagement
Best fitCoding and classification evalsLongitudinal, interval, and drift evals

What Each Path Strips From A Dental Chart

Most of what a dental chart contains is not identifying on its own. Tooth numbers, surfaces, CDT codes, periodontal probing depths, and restoration materials appear nowhere in the 18 identifiers. Accordingly, a dental charting model or a CDT code mapping pipeline can usually be evaluated on Safe Harbor data without losing signal.

Dates are where the two paths split. Safe Harbor reduces every visit to a year, which erases recare intervals, days between a crown prep and seat, and the pace of perio progression across successive charts.

Expert Determination usually allows a consistent per-patient date shift instead. Every date for a given patient moves by the same random offset of, say, minus 180 to plus 180 days, so intervals survive while the calendar anchor does not.

Safe Harbor keeps only the year of each visit, so the intervals between visits are lost. Expert Determination usually allows one random date offset per patient, which keeps the spacing between visits and hides the real dates.

Next, look at the fields that sit next to the clinical data. The ones below leak identity in almost every practice-management export we have seen from Dentrix, Eaglesoft, and Open Dental via its API:

  • Insurance subscriber and group numbers. These are health plan beneficiary numbers under Safe Harbor. They also appear inside the claim attachments and EOB text that get pulled into notes.
  • Guarantor and responsible-party records. The guarantor is often a parent or spouse, and a relative's identifiers fall under the rule too. Family account structures link records across the household.
  • Implant and device records. Implant lot and serial numbers are device identifiers. Registries and manufacturers can trace them back to a named patient.
  • Internal primary keys. PatNum and chart numbers are medical record numbers. Replace them with a keyed surrogate generated outside the source system.

Provider names and NPIs aren't patient identifiers under Safe Harbor. However, a single-provider practice in a rural three-digit ZIP can narrow the patient pool sharply, and an expert would flag that combination even when the checklist doesn't.

Radiograph Metadata Is Where Structured Scrubbing Stops

Imaging is where most dental de-identification pipelines break. Bitewings, periapicals, panoramics, cephalometrics, and CBCT volumes all carry identity in places a column-level scrubber never touches.

DICOM headers are the obvious layer. Tags such as PatientName (0010,0010), PatientID (0010,0020), PatientBirthDate (0010,0030), StudyDate (0008,0020), InstitutionName (0008,0080), AccessionNumber (0008,0050), and DeviceSerialNumber (0018,1000) must be removed or replaced, and Study and Series Instance UIDs must be remapped consistently. The DICOM standard's Basic Application Level Confidentiality Profile in PS3.15 Annex E is the reference action table.

Many intraoral sensor platforms also export JPEG, PNG, or proprietary formats instead of DICOM. Those files carry identity in EXIF blocks and in filenames such as SMITH_JOHN_20250312_BWX.jpg, and both need handling.

Clearing DICOM tags is not enough to de-identify a dental radiograph. Panoramic and ceph images often have the patient's name and date printed into the pixels, and a CBCT volume can be rendered into a recognizable face.

Burned-in annotation is the second layer. Panoramic and cephalometric units frequently stamp the patient name, date, and practice name directly into the image pixels, so the pipeline needs OCR detection and pixel masking before any image enters an eval set for AI radiograph analysis.

CBCT is the third layer. A full-skull volume can be surface-rendered into a recognizable face, which puts it squarely in "full-face photographs and any comparable images." This is why CBCT eval sets need defacing or field-of-view cropping, in addition to header cleanup.

Extraoral clinical photos follow the same logic. Intraoral photos showing only dentition generally fall outside the full-face category, but smile-design and orthodontic records often include full-face frontal and profile shots that must be excluded or masked.

Where Free-Text Clinical Notes Leak Identifiers

Structured fields fail loudly because a column is either there or gone. Free text fails quietly, and in our experience it is where most residual PHI in a "de-identified" dental dataset actually lives.

Clinical notes, treatment-plan comments, and referral letters leak identifiers in recurring patterns. The most common include:

  • Relatives and companions. "Pt's daughter Maria drove her today" names a relative, and relatives are covered identifiers. Pediatric notes name parents constantly.
  • Employers and occupations. "Works nights at the Route 9 distribution center" gives both an employer and a location. Rare occupations act as quasi-identifiers even without a named employer.
  • Dates in prose. "Wants crowns seated before her June 14 wedding" slips past date-field scrubbing. Relative references such as "last Thanksgiving" can also anchor a timeline.
  • Contact details typed inline. Front-desk staff paste callback numbers and email addresses into notes. Insurance correspondence pasted into notes carries subscriber IDs.
  • Rare events. A note that references a specific accident, a local news story, or a notable trauma can identify a patient uniquely. No named-entity model catches this reliably, which is why expert review matters.
  • Other providers. Referral letters carry the referring dentist's letterhead, address, and phone. They also frequently repeat the patient's full name and date of birth in the salutation block.

The fix is to run a PHI detector over every free-text field. Examples include Amazon Comprehend Medical's DetectPHI, a fine-tuned clinical NER model, or an LLM-based scrubber running inside your BAA boundary, as described in our write-up on running clinical AI on Bedrock. The dental NLP vocabulary matters here, since general-purpose detectors misfire on terms like tooth numbers and shade codes.

How Do You Know Your Scrubber Is Working?

You measure it the same way you measure any model: against a hand-annotated gold set. Build one from a few hundred notes, label every identifier span by category, and report recall per category, not a blended F1.

The math is unforgiving. At 99% token-level recall, a corpus of 50,000 notes averaging six identifiers each still leaves approximately 3,000 identifiers in place.

Check a PHI scrubber against a gold set of notes labeled by hand, and report recall for each identifier category. Recall matters more than precision, because every identifier the scrubber misses is a potential PHI disclosure.

Two design choices lower the residual risk. First, replace detected identifiers with realistic surrogates instead of [REDACTED] tags, so that a missed name blends into a sea of fake names. Second, apply the same scanner to eval outputs and traces, because a clinical note summarization model can copy a missed identifier straight into its summary.

Surrogates also keep the text realistic, and that matters for eval validity. A summarizer tested on notes riddled with bracket tags will score differently than it does on production text, which quietly corrupts the result you were trying to measure.

Choosing A Path For Your Eval Set

The right path depends on what the eval measures. Here's how the choice usually breaks down for clinical dental AI work:

  • Use Safe Harbor for point-in-time classification. CDT coding accuracy, tooth-level charting, and single-image radiograph findings rarely need exact dates or fine geography. Safe Harbor is faster, cheaper, and easy to explain to a customer's privacy officer.
  • Use Expert Determination for anything longitudinal. Recare adherence, perio progression, treatment-plan acceptance over time, and model drift monitoring all depend on intervals. Year-only dates destroy those signals.
  • Use Expert Determination for regional analysis. Fee-schedule and payer-mix evals need three-digit ZIP or finer regions. The expert can support that where population density allows it.
  • Run both when the corpus is mixed. Many teams maintain a Safe Harbor tier for broad engineering access and an expert-certified tier with tighter access for longitudinal work. The two tiers share a scrubber and differ in their generalization rules.

For every tier, the eval design itself should follow the discipline we describe in building clinical AI evals. De-identification decides what data you may use; the eval design decides whether the numbers mean anything.

Operational Controls That Hold Either Path Together

De-identification is a pipeline, and pipelines drift. The controls below are what keep a certified dataset certified.

  • Re-identification keys under 164.514(c). HIPAA allows a re-identification code only if it isn't derived from patient information and the mechanism isn't disclosed. Store the surrogate-to-source map in a KMS-encrypted table inside the covered boundary, with every read logged in CloudTrail.
  • Versioned scrubber and rules. Pin the NER model version, the ZIP-prefix table, and the date-shift policy for each dataset release. When any of them changes, rerun the gold-set recall check before the new version ships.
  • Output scanning. Scan eval traces, model outputs, and exported failure cases, not just inputs. Pair this with the defenses in our post on prompt injection in clinical AI, since an injected note can try to pull context from other records.
  • Know the limited data set alternative. A limited data set under 164.514(e) keeps dates, city, state, and five-digit ZIP. It is still PHI and requires a data use agreement, so it cannot go into a non-BAA eval surface.

All of these controls add up to one property: you can show an auditor what was removed, how you measured what remained, and who could reverse it. That evidence also feeds directly into the vendor reviews described in our SOC 2 guide for clinical AI vendors.

Scoping Your De-Identification Pipeline

Getting real dental records into an eval pipeline safely comes down to the same few decisions every time. Those decisions are which HIPAA path applies, how imaging and free text get scrubbed, and how residual risk is measured rather than assumed.

If you're standing up an eval harness for a dental charting, imaging, or note-summarization model and need realistic data without PHI leaving your BAA boundary, the NexV team builds and operates HIPAA-grade clinical AI pipelines for dental groups every week. Reach out for a working session. We'll map your source systems, choose the de-identification path each eval needs, and hand you a scrubber recall baseline you can take into your next security review.