Saying less is not seeing more
By Reza Motaghi
The newest AI model invented half as much on my radiographs. It also missed more disease. Which one is "safer"?
In September I gave the newest general AI model I could reach 128 of my own panoramic radiographs, blind, with the exact instructions its predecessor had been given in August. One board-certified oral and maxillofacial radiologist, every statement compared, finding by finding, against the reads my co-reader and I had signed, by a judge that could not see which model had written it. Hits and inventions were counted apart. There was no overall score, by design.
The new model invents less. Fewer than half as many invented findings. Of what it asserted, about one in three was wrong, down from about seven in ten. It counts the teeth better: more than nine in ten right, up from about eight in ten. Overall it found roughly the same share of what I read.
Then the two commonest findings in my reads. Bone loss: from about half of what I read down to under four in ten. Caries: from about one time in nine down to about one in eighteen. Its report is about a third of the length of the old one.
The day before, a model built for dentistry had answered the same task with yes or no and one stock paragraph. It said so little that there was little to invent, and what little it asserted was mostly wrong. It also found next to nothing.
Here is the shape of it. Any reader, human or machine, can look safer by saying less. Inventions fall, and so does the disease found, and a single safety number cannot tell the two apart. Hallucination rate, refusal rate, "hedges when unsure": each one rewards the quiet reader. The patient with the unreported caries does not care that the report was clean.
There is a name for this trade in the evaluation literature: precision against recall. Precision is the share of what the model said that was right. Recall is the share of what was there that the model found. A model moves along that line for free by talking less, and a scorecard that shows only one end of it will call the move an improvement. Radiology has known this for a century under other names: a reader who never calls a lesion has a perfect false-positive rate.
Text models have the mirror problem. A 2025 paper from OpenAI researchers argues that the usual benchmarks grade an answer right or wrong with no credit for saying "I don't know", so a model learns to guess (Kalai and colleagues, 2025). Both are the same lesson read from opposite ends: the model becomes whatever the scorecard rewards, and a scorecard with one number always has a cheap way to be gamed.
So my test measures both directions, per finding, and refuses to average them. It gates on inventions first, then asks whether the model still finds what the old one found, finding by finding, with a floor set before the run. By that rule the newest model does not replace the old one. It is a different reader, not a better one: quieter, more honest about the inventory, blinder on the two findings that fill most of my reports.
This is not a radiology rule. The same trade is on offer wherever a model is tuned to refuse more, hedge more, answer shorter. Quieter is sold as safer. Ask what it stopped seeing.
One caveat I owe the reader. Caries on a panoramic radiograph is hard for anyone. The overview image was never the film for it, and general models do poorly on it in the published tests too. The invention counts here are a blinded judge's and await my review of every flagged line, so they may move. The misses are counted against my own signed reads and stand.
What I would do differently: put the length of the answer on the scorecard from day one. A shorter answer moved every number, and I noticed after.
The check
The two-direction test, before you call a model safer:
- Count what it invented and what it missed as two numbers. Never one. An average lets the quiet model buy a pass.
- Put a floor under what it must find, per finding, before the run. The findings that fill your reports set the floor. The rare ones can wait.
- Measure how much it said. When the answer gets shorter, both numbers move for free. Fewer inventions with more misses is silence, not safety.
Sources
- Kalai AT, Nachum O, Vempala SS, Zhang E. Why Language Models Hallucinate. arXiv:2509.04664, 2025. "Hallucinations persist due to the way most evaluations are graded: language models are optimized to be good test-takers, and guessing when uncertain improves test performance."
- Liu Z, Ai QYH, Yeung AWK, Tanaka R, Nalley A, Hung KF. Performance of a Vision-Language Model in Detecting Common Dental Conditions on Panoramic Radiographs Using Different Tooth Numbering Systems. Diagnostics 2025;15(18):2315. A general model near perfect on implants, zero sensitivity for caries and periapical lesions on panoramic radiographs.
- Schwendicke F, Tzschoppe M, Paris S. Radiographic caries detection: a systematic review and meta-analysis. J Dent 2015;43(8):924-933. Pooled sensitivity for any carious lesion on radiographs between 0.24 and 0.42.
- The test itself: 128 of my own cases, blind, two general models given the same instructions six weeks apart, every statement checked finding by finding against the two-reader reads. The models' counts are in the private record; the answer key's own counts stay there.