Skip to content
One-Person Startup12 min read

What an AI model assumes when a patient fact is missing

By Reza Motaghi

A row of identical adult figures drawn in a single white line on black, one small child among them casting a long adult-sized shadow, a minimalist illustration about the patient an AI model writes when it was never told who is there
On this page

Radiology has had a default patient since 1975. The standard that defined him, ICRP Publication 23, calls him Reference Man.1 He is between twenty and thirty years old, weighs 70 kg, stands 170 cm, and is "a Caucasian and is a Western European or North American in habitat and custom". He exists so that radiation doses could be calculated once and reused. For decades everyone else, women and children included, was scaled from him. The standard did not get a Reference Female until 2002.2 The two of them became the computational phantoms that dose calculations run on today only in 2009.

Every AI model has a default patient too. Nobody wrote him down. He is whoever appears most often in the training data, and he walks into the room the moment a fact is missing. Outside medicine he has other names: the default codebase, the default contract, the default customer. The failure is the same. A confident answer about someone else's case, and only the person who knows the real case can tell.

What the default patient is

The default patient is the case an AI model writes when a fact was never given. Not the case in front of it. The most common case it has seen, delivered with the confidence of a case it was told. He is not an average. An average of the children and the adults in a dataset would be a teenager. He is the majority case, the one that both the training data and the scoring reward. He arrives without a hedge because he was never a guess to the model. He was the shape of the world.

Once you have a name for him you start seeing him everywhere, which is the point of naming him.

The patient in the report did not exist

A child's radiograph, baby teeth still in the mouth. Between about six and twelve years of age a child carries both sets of teeth at once.3 The primary teeth are on their way out and the permanent teeth on their way in. The two sets are even numbered differently, twenty primary teeth lettered A to T, thirty-two permanent teeth numbered 1 to 32. Any dentist reads that image as a child within a second.

The report that came back charted a full set of thirty-two permanent teeth and found nothing unusual. Fluent, complete, confident. Not one word flagged, not one hedge. The frontier model that wrote it had never been told this was a child, so it wrote the patient it sees most often. An adult. It did not get the case wrong. It answered a different case, and it answered it well.

That is the tell worth learning. The failure did not look like a failure. It looked like a good report.

Where the default patient comes from

He comes from the data first. Across 181 public medical imaging datasets, children are under 1 percent of the patients whose age is recorded at all.4 Of the AI-enabled medical devices the FDA has cleared, only 4.4 percent carry a pediatric age range, and nearly 60 percent carry no age information.5 A model that has barely seen a child does not know that it has not.

The result shows in the reading. When adult-trained imaging AI is put in front of children without adaptation, performance drops in eighteen of twenty studies, and the youngest children lose the most.6 For lung nodules, sensitivity that ran up to 100 percent in adults fell as low as 26 in children.

He comes from the scoring second, and this is the part that generalizes past medicine. Models are evaluated so that a confident guess scores better than "I do not know".7 In the human preference data used to shape them, answers that hedge are rated lower than answers that do not.8 Frontier models asked to state a confidence cluster between 80 and 100 percent whether or not they are right.9 Put those together and the behavior is not a bug. A model that fills the gap with the majority case and sounds sure is doing exactly what it was rewarded for.

So he shows up in every field. Asked to write clinical vignettes with no demographics given, one frontier model produced patients that stereotype the condition.10 Coding assistants reach for whichever API surface was most common in their training data, deprecated calls included.11 They recommend packages that do not exist about one time in five.12 General models get real legal questions wrong more often than they get them right.13 Every one of those answers read like a good one.

Medicine has met the default patient before

This is not a new failure. It is an old one running at a new speed. In 1968 the pediatrician Harry Shirkey called children "therapeutic orphans".14 Drugs were tested on adults, and children got a scaled-down dose by formula. Clark's rule took the adult dose, multiplied by the child's weight in pounds and divided by 150. Thirty years later the United States Senate recorded that under 20 percent of prescription drugs were labelled for pediatric use.15 It took two acts of Congress, in 2002 and 2003, to change it.

Medicine needed decades and legislation to learn that scaling from the default is not the same as knowing the patient. AI models are re-running that history in months, on every input at once. There is no act of Congress coming for your codebase or your contract. The check has to be yours.

Four checks before you sign

The checks are short because they run on every off-average input. They are honest about their own limits because a skeptical reader deserves that. Two of them I changed after reading the evidence against them.

1. Name the difference before you ask, then make the model redo what it assumed. State up front what makes this case not the majority case. A child, not an adult. A legacy module, not the main branch. A customer in an edge market, not the median one. Necessary, not sufficient. Even when the fact is stated, the pull toward the default can survive in the parts of the answer written before the fact was used. So after the answer, ask for one more pass: revise anything that assumed the default. On the child's radiograph the difference is one line, "mixed dentition, primary teeth present," and it changes the whole chart.

2. Ask for the assumptions, then check the assumptions, not the prose. Ask the model to list every fact it assumed that you did not give: age, version, jurisdiction, region. Then verify those three or four lines instead of rereading the paragraph. One warning here. The list is a probe, not a confession. In one lab's own tests, reasoning models mentioned the hint they had actually relied on in well under half of cases, and less often on harder problems.16 So treat every assumption it names as a lead to check. Treat the ones it does not name as still open.

3. Read the missing hedge as the flag. A clean, complete answer to a messy input is the warning sign, because the model was trained to sound sure. Ask what it was unsure about. A model that was unsure about nothing is the one to distrust. This check has an off-ramp. On a routine, well-covered case, forcing uncertainty out of a model costs accuracy and produces refusals it did not need to make.17 Run this on the input that is not the average, not on every input.

4. Once, on purpose, hand it an off-average case. Take a case you know cold that sits outside the majority, and give it to the model with the distinguishing fact left out. See whether it notices. If it reads the child as an adult, you have learned what its silence is worth on every case you cannot check yourself. This is the cheapest evaluation there is. It is also the one most people never run, because the fluent answers on the easy cases feel like enough. Run it again after every model update, since an update can change behavior your workflow depends on without changing the tone of the answers.

The default patient, four checks. One, name the difference before you ask, then make the model redo what it assumed. Two, ask for the assumptions and check those, not the prose. Three, read the missing hedge as the flag. Four, once, on purpose, hand it an off-average case.

Two things the checks are not. They are not an argument that AI cannot read children. Built or even re-thresholded for children, imaging AI reads them well. A pediatric fracture model held its sensitivity across ages and regions in a multicenter evaluation.18 An adult chest X-ray tool reused on pediatric films did respectably and improved further with a pediatric threshold.19 The problem was never the child. It was the fact that nobody stated. And the checks are not a prompt trick. They are the discipline a reading room applies to any report that arrives without a history. Before you sign it, find out who it is about.

What I sign for

The average patient is the one who never shows up. Twenty years of reading radiographs taught me that. Writing training examples from my own reads taught me the other half. The majority case is exactly what a model reaches for first, and it reaches for it quietly. Evaluating frontier vision models on my own cases, blind, is now my main research work. The default patient is the failure I look for before any other, because it is the one that leaves no mark on the page.

A radiologist signs for what was seen. A model reports what was likely. It is the same line I built into my own CBCT viewer, where the agent moves the camera and the human reads. Both can be right on the same image. The day they are not, only one of them knows who the patient was. Run the fourth check on the next model you rely on. If you build or evaluate models that look at images and want a radiologist who builds on your side of the table, the contact page reaches me directly.

References

Footnotes

  1. ICRP Publication 23, Report of the Task Group on Reference Man. International Commission on Radiological Protection, 1975. Reference Man is defined as "between 20-30 years of age, weighing 70 kg, is 170 cm in height ... He is a Caucasian and is a Western European or North American in habitat and custom." https://www.icrp.org/publication.asp?id=ICRP+Publication+23

  2. ICRP Publication 89, Basic anatomical and physiological data for use in radiological protection: reference values, 2002 (the Reference Female). ICRP Publication 110, Adult reference computational phantoms, 2009. https://www.icrp.org/publication.asp?id=ICRP+Publication+89 and https://www.icrp.org/publication.asp?id=ICRP+Publication+110

  3. American Academy of Pediatric Dentistry, guidance on the mixed dentition period, about six to twelve years. Tooth numbering: ADA Universal Numbering System, primary teeth A to T, permanent teeth 1 to 32.

  4. Zamora Hua et al. Pediatric representation in public medical imaging datasets. medRxiv, 2025. 181 datasets, children under 1 percent of patients where age is documented. https://pmc.ncbi.nlm.nih.gov/articles/PMC12259205/

  5. Zapotoczny et al. Pediatric labelling of FDA-cleared AI-enabled medical devices. JAMA Network Open, 2026. 952 devices, 42 with a pediatric age range, nearly 60 percent with no age information. https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2846735

  6. Laborie et al. Performance of adult-trained artificial intelligence models in paediatric imaging, a scoping review. European Radiology, 2026. Twenty studies, all but two with reduced performance, greatest deficits at two years and under. Lung nodule sensitivity 68 to 100 percent in adults against 26 to 68 in children. https://doi.org/10.1007/s00330-026-12354-5

  7. Kalai A, Nachum O, Vempala S, Zhang E. Why language models hallucinate. OpenAI, 2025. https://arxiv.org/abs/2509.04664

  8. Zhou K, Hwang JD, Ren X, Sap M. Relying on the unreliable: the impact of language models' reluctance to express uncertainty. Association for Computational Linguistics, 2024. https://arxiv.org/abs/2401.06730

  9. Chhikara P. Mind the confidence gap: overconfidence, calibration, and distractor effects in large language models. Transactions on Machine Learning Research, 2025. https://arxiv.org/abs/2502.11028

  10. Zack T et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. Lancet Digital Health, 2024. https://pubmed.ncbi.nlm.nih.gov/38123252/

  11. Wang C et al. LLMs meet library evolution: evaluating deprecated API usage in LLM-based code completion. International Conference on Software Engineering, 2025. https://arxiv.org/abs/2406.09834

  12. Spracklen J et al. We have a package for you! A comprehensive analysis of package hallucinations by code generating LLMs. USENIX Security, 2025. https://arxiv.org/abs/2406.10279

  13. Dahl M, Magesh V, Suzgun M, Ho DE. Large legal fictions: profiling legal hallucinations in large language models. Journal of Legal Analysis, 2024. Hallucination rates of 58 to 88 percent on verifiable legal questions. https://doi.org/10.1093/jla/laae003

  14. Shirkey H. Therapeutic orphans. Journal of Pediatrics, 1968. Clark's rule: adult dose multiplied by the child's weight in pounds, divided by 150. https://doi.org/10.1016/S0022-3476(68)80414-7

  15. United States Senate Report 105-43, 1997, cited in Institute of Medicine, Safe and Effective Medicines for Children, National Academies Press, 2012. The Best Pharmaceuticals for Children Act, 2002, and the Pediatric Research Equity Act, 2003. https://www.ncbi.nlm.nih.gov/books/NBK202040/

  16. Anthropic. Reasoning models don't always say what they think. 2025. Models mentioned the hint they relied on in about 25 to 39 percent of cases, less on harder tasks. https://www.anthropic.com/research/reasoning-models-dont-say-think

  17. Wen B et al. Know your limits: a survey of abstention in large language models. Transactions of the Association for Computational Linguistics, 2025. Over-abstention as a first-order cost. https://arxiv.org/abs/2407.18418

  18. Multicenter evaluation of a deep learning model for pediatric fracture detection. Academic Radiology, 2026. Sensitivity 0.96 overall, above 0.93 across age and regional subgroups. https://pubmed.ncbi.nlm.nih.gov/41320594/

  19. Deep learning for pediatric chest x-ray diagnosis: repurposing a commercial tool developed for adults. PLOS ONE, 2025. 958 pediatric radiographs, sensitivity improved with a pediatric-specific threshold. https://doi.org/10.1371/journal.pone.0328295