Four pillars. One practice.
The four pillars
Four constituencies in pediatric care, each with a different stake and a different way of being failed.
Most pediatric decisions get made at home, at night, by someone who is not a clinician and has no one to ask. That is where the worry actually lives, and it is the least served part of the system.
Fails as: an assistant that answers confidently and sends a sick child back to bed.
The people carrying the work. Tooling should give time back to the encounter rather than extract more throughput from the person doing it.
Fails as: efficiency gains that quietly become the new expected baseline.
The people who will practice after us. Using these tools well is becoming part of clinical competence, and very little of it is being taught deliberately.
Fails as: teaching the tool instead of the judgment.
The pillar that keeps the other three accountable. If we cannot measure whether something helps, we are guessing with other people's children.
Fails as: reporting the number that flatters the product.
A pediatric DPC AI scribe. Browser-first, bring-your-own API keys, no patient data on our servers.
Try BRTLB →HIPAA-compliant AI assistant framework that turns Spruce Health DMs into a secure command channel for clinicians.
Email for access →Turns clinical text into a safe, paste-ready prompt for an outside AI. De-identification masks identifiers and takes the medicine with it; decon strips the identifiers and keeps the clinical facts. Runs on your own machine.
The pediatric intelligence layer — clinical AI tuned for the patterns, dosing, growth curves, and red flags adult systems miss.
Tour Northstar →Strategy, product posture, and clinical-AI evaluation work with health-tech teams.
A pediatric medical-education ecosystem: synthetic patients, voice encounters with error injection, a teaching EMR, and a Socratic AI tutor.
A practical curriculum on AI in clinical practice — for clinicians who want to use these tools well, not just hear about them.
Open AI101 →Independent benchmarks and evaluation methodology.
Scenarios start the way visits do: an undifferentiated complaint in a parent's words, across five age bands where the same fever means different things. Scoring counts what a model left out as heavily as what it got wrong, and penalizes ordering everything. A held-out set stays private.
One routine ear infection with two facts left blank. 41% of frontier-model answers invented the missing history; six of seven error types showed up, and bias was second rarest. Dedicated clinical tools fixed the stale dose and mostly asked instead; one kept fabricating with citations attached. One added sentence cut invention from 43% to 11%.
The commercial clinical AI tools scored on what a clinician actually opens them for: is the citation real, is the guidance current, does it refuse a false premise, does it ask before it assumes. Eight tools have run the ear-infection chart in the open packet. The broader battery, with recency traps and fabricated-guideline traps, is still being built.
Read the methodology →Featured writing
Essays on what clinical AI actually does when you test it, and what to build instead.
Every error the models made has a name in the human literature. An honest comparison of two differently fallible systems, and why adding a doctor to the loop is not automatically safer.
Ten frontier models, one synthetic ear infection, 140 traces. Where the invented facts came from, what identity changed, and the one sentence that cut fabrication by three quarters. Findings and packet →
Five different things get called de-identification. In pediatrics the fields that identify a patient are often the fields that make the answer correct, so deletion breaks the medicine. The dual-gate evaluation behind Clinician Decon.
Clinicians have always worked from a pyramid of evidence, gut at the base and meta-analyses at the top. AI collapses retrieval to seconds and hides source quality, so the thing to teach is the epistemology, not the answer.
Disaggregate the Nature Medicine paper and the gap is readability. Accuracy and safety tied, and the two things a reference tool exists for, citations and recency, were never scored.
Prompting is not coding and not writing either. It sits between disciplines product teams already have, and a prompt for a pediatric triage system draws on clinical training as much as on either.
Every physician can describe the best assistant they ever worked with, and none of it is about credentials. Not a tool, a teammate: one that knows when to interrupt, anticipates the next question, and can be handed something.
The most common procedure in medicine is the interview. Triage bots and symptom checkers are being built on it too, and if AI is going to be trusted with health it has to learn how that conversation heals.
A credible medical model small enough to run beside the EHR, open enough to shape, and yours to run without sending patient data anywhere. Discharge instructions in Thai, generated on a laptop.
Still building
Public and private commits across every project on this page.