Clinician Decon
In pediatrics, the fields that identify a patient are frequently the same fields that make the answer correct. Deleting them protects the record and breaks the medicine.
Clinicians are already pasting clinical text into general AI tools. That happens today, in every practice, with or without a policy about it. The useful question is not whether to stop it — it is what a physician should be able to paste, and what a tool should do to the text first.
The obvious answer is to strip anything that looks like an identifier. That answer is wrong, and it is wrong in a way that is hard to see, because the output looks safer while the clinical question quietly stops being answerable.
"De-identification" gets used for at least five distinct operations with different legal standing, different threat models, and different failure modes. Conflating them is the root of most bad tooling in this space.
| Operation | What it actually asserts | Where it breaks for prompts |
|---|---|---|
| Removing obvious identifiers | Nothing formal. A pattern sweep for names, numbers, and dates. | No standard behind it, and it silently destroys clinical fields that happen to look like identifiers. |
| HIPAA Safe Harbor | A regulatory determination: the eighteen enumerated identifier categories are gone, including all elements of dates and ages over 89. | It is a document-release standard. Applied to a prompt it removes exact dates and age granularity wholesale — which in pediatrics is the clinical variable, not metadata. |
| Expert Determination | A qualified statistician certifies the re-identification risk is very small for a defined data set and recipient. | It is a determination about a corpus and a context, not a per-message operation. There is no such thing as expert-determining an ad-hoc paste at 2am. |
| Pseudonymization | Identifiers are replaced with consistent surrogates and a re-linkage key exists somewhere. | Because the key exists, the data is generally still regulated. Consistent surrogates also leak across messages — the same token recurring is itself a signal. |
| PHI minimization before a downstream send | Nothing about legal status. Only: less identity information left the machine than went in, and a human reviewed what left. | This is the honest framing for prompts — and the one nobody has built good tooling for. |
Clinician Decon does the fifth thing. It does not perform, claim, or approximate the first four. That scoping is not modesty; it is what makes the design tractable.
The dangerous assumption is that identity information and clinical information are separable — that you can pull one out and leave the other intact. In clinical text they are frequently the same tokens.
[DATE] destroys the chronology while leaving the sentence grammatical, which means the model answers anyway, without the thing that determined the answer.So a naive scrubber fails in both directions at once. It removes clinical signal it should have kept, and it retains re-identifying context it should have removed — because the first is easy to pattern-match and the second is not.
Treat this as clinical decontextualization with semantic preservation, not regex-and-delete. The transformed case has to retain clinically meaningful temporal relationships and attributes while minimizing unnecessary identity information. Those are two objectives, and a system optimizing only the second will reliably score well and answer worse.
Semantic preservation needs an operational test, or it collapses into "keep whatever seems useful." The rule the tool implements is a conjunction on both sides.
The third keep-condition is the one that does the real work. It converts rather than deletes: a birth date becomes an age band, an exact lab becomes a severity category, a specific location becomes a region — when the question does not need precision. When it does, the value stays.
Every example below is synthetic. No real patient data appears on this page.
A birth date is an identifier. The age is the answer.
Noah Rivera DOB 3/15/2013 MRN LP-2024-08432 needs catch-up vaccines before school starts Monday.
[NAME] DOB [DATE] MRN [ID] needs catch-up vaccines before school starts [DATE].
[NAME] adolescent [MRN] needs catch-up vaccines before school starts [DATE].
The exact birth date is identifying and unnecessary. The age band is clinically essential, because the schedule depends on it. The difference is not that one redacts more — it is that one converts the identifier into the signal the clinician actually needed.
Sometimes the exact number has to survive.
Maya Thompson is 28 kg and needs epinephrine autoinjector dose after peanut anaphylaxis.
[NAME] is [NUMBER] kg and needs epinephrine autoinjector dose after peanut anaphylaxis.
[NAME] is 28 kg and needs epinephrine autoinjector dose after peanut anaphylaxis.
Weight is quasi-identifying in the abstract and decisive here — 28 kg sits near the autoinjector dosing threshold. Redacting it produces output that looks safer and answers worse. This is the case that breaks any policy expressed purely as a list of field types.
The same token can be a patient and a diagnosis.
Referral for Addison Brooks. Addison disease on fludrocortisone, vomiting; stress-dose steroid?
Referral for [NAME]. [NAME] disease on fludrocortisone, vomiting; stress-dose steroid?
Referral for [NAME]. Addison disease on fludrocortisone, vomiting; stress-dose steroid?
A name-matching pass sees the string twice and masks both. One of them was the diagnosis, and the question was about adrenal crisis. Eponyms — Addison, Bell, Graves, Still, Crohn, Hodgkin — make this a systematic failure, not an edge case.
Uniqueness without an identifier.
Only HLH patient in Room 712 today: ferritin 18000, platelets 22, fibrinogen 92. Does this meet treatment criteria?
Only [CONDITION] patient in [LOCATION] [DATE]: ferritin [NUMBER], platelets [NUMBER], fibrinogen [NUMBER]. Does this meet criteria?
Rare HLH case: ferritin 18000, platelets 22, fibrinogen 92. Does this meet treatment criteria?
Inverted on both axes. The stripped version keeps the re-identifying part — "only patient with this condition, this unit, this day" narrows to one person — and destroys the labs, which are literally the diagnostic criteria being asked about. The preserved version drops the uniqueness context and keeps the numbers.
Two objectives means two gates, and a system is only shippable if it clears both on the same case. Optimizing one is easy and produces a tool that is dangerous in a way its own metrics cannot see.
| Term | Definition used in scoring |
|---|---|
| PHI-leaked output | At least one labeled forbidden string survived into the destination output. |
| Unsafe copy-allowed | A leaked output that the pipeline nonetheless marked as safe to hand off. The worst category — a wrong answer delivered with confidence. |
| Missing critical fact | A fact labeled in advance as clinically required for that case did not survive the transform. |
| Handoff-usable | No labeled leak, and no missing labeled critical fact, and copy allowed. The only outcome that counts as a pass. |
Labeling the required facts per case, before running anything, is the part that makes the second objective measurable. Without it "clinical usefulness" stays a vibe and every scrubber looks fine.
The release corpus is now 3,587 source cases rendered to 10,761 destination outputs (chat, a second chat product, and web search), built up from two generated 500-case development suites (usability and adversarial), a 2,000-case persona regression, several small authored blind-spot suites aimed at known weaknesses, a clinician-authored seed set, a real-world-shaped adversarial set, and the structured-record and cross-turn sets described below. Every suite is generated, authored, or synthetic. No real patient data is used at any stage.
Each stage below looked like success at the time. The re-test is what changed the architecture. All figures are on synthetic fixtures.
| Stage | What the evidence showed | What it forced |
|---|---|---|
| Semantic rewrite partner engagement |
Zero true leaks across 132 adversarial queries, with one validator false positive. | Semantic transformation preserves intent better than literal masking — but it leaned on a cloud model under a BAA, and over-minimized a full clinical prompt. |
| Local NER baseline | 1 of 18 identifiers leaked in the first 10-case comparison; a deterministic prefilter closed the record-number gap. | Combine high-confidence patterns with contextual NER. Neither is sufficient alone. |
| Distribution challenge | Clean synthetic templates: zero leaks. The combined 1,132-query set: 79 of 2,526 expected terms missed. | Conversational, multilingual, multi-patient, family, URL, practice, and measurement cases matter far more than tidy templates. Passing on clean data predicts nothing. |
| Replaying the shipped engine | The polished rules-only build leaked 1,067 terms toward a chat destination and 915 toward a web destination — and marked every leaking case low-risk and copyable. | Bind release evidence to the engine that actually ships, and fail closed. Interface confidence is not pipeline evidence. |
| The utility gate | 415 of 1,500 outputs lost a clinically critical fact — concentrated in pediatric dosing weights and criteria-relevant labs. | The thesis of this page, with a number on it: privacy-only optimization produces clinical harm. Label what must be preserved and measure it. |
| Rules + NER, tuned | A broad first gate hit zero leaks but lost required facts in 174 of 3,000 outputs. After targeted fixes, all 3,000 cleared both gates. | What works is clinical token shields, semantic transforms, destination-specific policies, residual validation, and an explicit engine — not more redaction. |
| Browser prototype earlier internal build |
The architecture called for local transform, visible correction, destination-safe send, and fail-loud behavior. The archived demo was still doing client-side regex. | The clinician interaction was right; the engineering wasn't. Shipping needs a validated browser model, a parity gate, server-side defense, and revalidation after any user correction. |
| The stale artifact August 2026 |
Repository source masked a first name inside a note parenthetical on both engines. The packaged Mac app on the same machine, built from a pre-hardening snapshot, served it unmasked — marked low-risk and copyable — while every source-level gate stayed green. | Validation had only ever run against source. Now a smoke gate posts ten known-adversarial cases at the running build itself, and a build is not a release until the artifact passes. |
| Residual fragments September 2026 |
A red team of twelve structured-record shapes (FHIR, HL7, JSON, OCR) failed six of them: 17 of 36 outputs kept identifier material, and all 17 allowed copy. The evaluator was searching for the exact source string, so [MRN]DEFGHIJK counted as clean. |
Score remnants, not strings — an output fails if a residual token or the visible fragments around a placeholder preserve half the original value. Authoritative structured-field spans, propagation of confident detections, and a partial-remnant guard. Now 36 of 36, and the expanded corpus 10,761 of 10,761. |
Stage five is the one worth sitting with. The system was getting safer and less useful at the same time, and only one of those was instrumented. The last two rows are the same lesson twice more: the gate was green because it was measuring the wrong artifact, then the wrong string.
If the model layer is part of the safety argument, then the system has to know what to do when that layer is missing. Running the same 72 outputs through three engine states makes the design visible.
| Engine state | PHI leaked | Leaked and copy allowed | Handoff-usable |
|---|---|---|---|
| Rules only, model absent | 3 / 72 | 3 | 69 |
| Model unavailable, gate engaged | 3 / 72 | 0 | 0 |
| Rules + NER, model present | 0 / 72 | 0 | 72 |
The middle row is the whole point. The same three leaks are present, but nothing is marked copyable — the tool refuses to hand off rather than degrading quietly into the top row, which is the genuinely dangerous configuration: mostly working, confidently wrong three times out of seventy-two, and telling the clinician it was fine.
That middle row is now the shipped default. The app, the CLI, and the validation command all run rules plus the local model; if the model is missing they report high risk and block copy rather than fall back to rules alone. The rules-only engine still exists for regression work and has to be asked for by name.
This comparison is favorable by construction and should be read that way. The rubric is ours. The cases are ours. The comparator is a conventional de-identification tool being scored on a clinician-prompt task it was not built for, and it required compatibility work to run in our environment. An earlier adversarial add-on was used to harden our own pipeline before the final run. It is a directional result about task-fit, not a neutral benchmark.
With that said: across 1,000 shared source cases scored on the dual gate, the external de-identification baseline cleared the safety gate on 75.9% and the usability gate on 39.9%. The task-aware pipeline cleared both on 100%.
The usability number is the informative one. A conventional de-ID tool doing exactly what it was designed to do renders roughly six out of ten clinical prompts unanswerable — not because it failed, but because document release and prompt construction are different jobs and only one of them was its job.
The cleanest test of the thesis arrived in September 2026, when a newer PII-masking model was published with results that beat the local model this tool ships on the publisher's own benchmark, especially on identifiers that recur across conversation turns. The obvious move was to swap it in. So both models were run through the identical pipeline — same deterministic rules, same span composition, same destination rendering, same evaluator — and scored on the product task rather than on raw detection.
On the frozen release corpus of 3,171 source cases, the shipping model kept a perfect product score. The candidate erased clinical facts in 221 source cases and leaked an EHR wrapper field in one. On a new 260-case cross-turn stress set built to favor it, the candidate did cut recurring-identifier leaks from 10 cases to 3 — and in the same run made 61 cases unusable and damaged 20 clean sensitive-clinical controls that contained no identifiers at all.
A stronger detector, measured on detection, produced a worse clinical prompt, measured on the prompt. Nothing about the pipeline changed except the model. The second gate is what made that visible.
The candidate stays in the repository as an experimental adapter with the comparison harness. If multi-turn masking becomes a product requirement it is worth revisiting as a secondary guard, after label-specific clinical calibration and a much larger clinician-reviewed clean-negative set. The raw comparison was recorded before the hardening in the table above, and that hardening does not retroactively change it.
Direct handles: names including relatives' and providers'; chart, claim, accession, and device identifiers; phone numbers, emails, URLs, IP addresses; street addresses and ZIP-level geography; schools, camps, facilities, rooms, units, beds, floors; exact dates and birth dates. It also strips prompt-injection text, including instructions that attempt to talk the pipeline into preserving identifiers — a paste-driven tool has to assume its input is adversarial, because sometimes the input is a forwarded message.
What it protects: age bands derived from birth dates; decision-relevant labs when triage or criteria depend on them; exact weights where the answer is weight-based dosing; relationship roles such as parent reports patient; clinical eponyms and diagnoses when they are not functioning as identifiers; and the temporal relationships between events even when the absolute dates are gone.
This is not a legal de-identification certification tool, and running text through it is not a HIPAA compliance determination. It is a local PHI-minimizing prompt builder with a human review step that is not optional. The review step is the safety mechanism. The pipeline exists to make that review fast enough that it actually happens, rather than to replace it.
It is a research release. There is no formal threat model behind it yet, no independent security review, and no production hardening. Evaluation has run on synthetic fixtures and adversarial sets built specifically to break it — which is the right way to develop it, and is not the same as proving it safe in a live clinic. The Mac build is unsigned and not notarized. The hardening in the table closed the failure families that were found; it does not establish completeness on real-world text the fixtures never represented, and the residual-fragment episode is a reminder that the next family is found by the next red team, not by the last green run.
There is also a real tension the design does not resolve: every fact preserved for clinical utility is a fact retained for re-identification. The keep-rule narrows that surface but does not eliminate it, and a system that never made that trade would be one that always returns an unanswerable prompt. Anyone claiming to have solved this rather than traded against it is measuring only one of the two objectives.
The package, the fixture library, and the evaluation notes are public: github.com/dochobbs/clinician-decon. The web app runs from a checkout on 127.0.0.1 with no API key. There is also a self-contained Mac app — a native window around that same local server, with the Python runtime and the model files bundled so a clinician never sees a terminal — built from the repository with one script. It is unsigned, so macOS will warn on first open. If you're building something adjacent and want to compare notes, doc@4bv.ai.