Skip to content
Lab Notes6 min read

It counts the teeth and misses the disease

By Reza Motaghi

The model counted every implant in 128 of my panoramic radiographs. It never once called the jaw joint.

Blind, on 128 of my own panoramic radiographs, I had the strongest general AI model I could reach fill a fixed form: what is in the mouth, what is wrong with it, and whether the image is fit to read. One board-certified oral and maxillofacial radiologist, every answer checked against the reads my co-reader and I had signed.

The parts list, it got right. Implants: the exact number on all 128. Root fillings: matched my reads 97 times in a hundred. Fixed prostheses and edentulous spans most of the time, the count of teeth within one tooth on two cases in three. This is the part of a radiograph anyone can count. Earlier this summer the previous model, writing a full report on the same cases, got that list wrong hundreds of times. That is a note of its own. The newest one has fixed the count.

It has not fixed the reading. Caries: it called it present on three cases in the whole set, in a set where, by my reads, most cases carry it. On the rest it answered "uncertain". Periapical lesions, the same story. And the remodeling of the jaw joint that I report on about half of my panoramic radiographs: never once called present. Absent or unsure, on every one of the 128.

Then I gave it more time to think. Effort, in these tools, is how long the model reasons before it answers. At eighteen times the cost it found a handful more lesions, ranked the cases worse, and still did not call a single joint. More thinking did not buy more seeing.

A 2025 study put another general model through the same kind of test on panoramic radiographs and found the same split: implants near perfect, caries and periapical lesions at zero, which its authors put down to bright things on the image being easier to see than dark ones (Diagnostics, 2025).

Here is the shape of it. A general model is now reliable on the inventory and unreliable on the judgment, and it fails the judgment quietly: not with a wrong answer but with "uncertain", or with a calm "nothing to see" on the structure it does not know how to read. That is exactly where a fluent answer is most dangerous, because the half it gets right earns trust for the half it never did.

This is not a radiology rule. Ask an AI to list the clauses in a contract and it will, correctly. Ask which clause loses you the case, and read that answer against your own. The parts list is solved. The call is not, and the model will not tell you which of the two it is handing you.

One caveat I owe the reader. Caries on a panoramic radiograph is hard for anyone. It is an overview image, radiographic caries detection has low sensitivity in the meta-analyses even on the films made for it, and on close-up films a general model does much better. The joint is harder too: trained readers catch condylar flattening on this image about two times in three. Harder is not never. The model was asked about it by name and did not catch it once. And I asked for a form, not a report, with "uncertain" as an allowed answer. It used that answer the way a careful reader would. A careful reader who never calls the joint is still not reading the whole image.

For builders, one mechanical finding. The tool I used quietly shrinks any image wider than about two thousand pixels, and a panoramic radiograph is wider than that. Check the size your model actually receives before you judge what it saw. I sent each image in pieces, at full resolution.

The check

The split check, before you act on an AI read of anything:

  1. Ask for the inventory and the judgment as two answers. What is there, then what it means. One fluent paragraph hides which half you are reading.
  2. Check the inventory once, the judgment every time. Against your own read. The part that changes the decision is the part the model hedges.
  3. Read what it did not say. List what you always look at and tick it against the answer. The miss you never see is in the silence, not the prose.

Sources

  • Liu Z, Ai QYH, Yeung AWK, Tanaka R, Nalley A, Hung KF. Performance of a Vision-Language Model in Detecting Common Dental Conditions on Panoramic Radiographs Using Different Tooth Numbering Systems. Diagnostics 2025;15(18):2315. A general model at 98.8 percent balanced accuracy for implants, zero sensitivity for caries and periapical lesions, the authors attributing the split to radiopaque conditions being easier to see than radiolucent ones.
  • Camlet A, Kusiak A, Ossowska A, Świetlik D. Advances in Periodontal Diagnostics: Application of MultiModal Language Models in Visual Interpretation of Panoramic Radiographs. Diagnostics 2025;15(15):1851. General models in substantial agreement with clinicians on tooth counts, their bone-height measurements not usable.
  • Schwendicke F, Tzschoppe M, Paris S. Radiographic caries detection: a systematic review and meta-analysis. J Dent 2015;43(8):924-933. Pooled sensitivity for any carious lesion on radiographs between 0.24 and 0.42.
  • Pornprasertsuk-Damrongsri S, Vachmanus S, Papasratorn D, Kitisubkanchana J, Chaikantha S, Arayasantiparb R, Mongkolwat P. Clinical application of deep learning for enhanced multistage caries detection in panoramic radiographs. Sci Rep 2025;15:33491. "The detection of dental caries is typically overlooked on panoramic radiographs."
  • Stera G, Giusti M, Magnini A, Calistri L, Izzetti R, Nardi C. Diagnostic accuracy of periapical radiography and panoramic radiography in the detection of apical periodontitis: a systematic review and meta-analysis. Radiol Med 2024;129(11):1682-1695. Panoramic accuracy for apical lesions 66 percent.
  • Im YG, Lee JS, Park JI, Lim HS, Kim BG, Kim JH. Diagnostic accuracy and reliability of panoramic temporomandibular joint radiography to detect bony lesions in patients with TMJ osteoarthritis. J Dent Sci 2019. Sensitivity for condylar flattening under 67 percent, for osteophytes under 30 percent, fair to moderate agreement between readers.
  • Hatipoğlu Ö, Baytar M, Pertek Hatipoğlu F. Diagnostic and localization performance of multimodal large language models in the interpretation of dental periapical radiographs. BMC Oral Health 2026;26:1428. On close-up films, caries accuracy up to 73 percent for the best general model.
  • The vendors' own documentation on image size (opened 23/09/26): images above a model's long-edge limit are downscaled before the model sees them, without an error unless the caller asks for one; a 2,000-pixel limit applies in requests carrying many images. A published panoramic dataset stores its images at 2041 by 1024 pixels (Sensors 2021;21(9):3110).
  • The test itself: 128 of my own cases, blind, a fixed form checked against the two-reader reads. The model's answers are counted in the private record; the answer key's own counts stay there.

About the author

Reza Motaghi is an oral and maxillofacial radiologist and Chief Innovation Officer who reads every day and builds, evaluates, and trains the imaging AI for it. He built CBCTScope, the first CBCT viewer with native AI-agent control, and writes about what a model did on a real case and what it did not.

The newsletter

Get the essays as they’re written.

Imaging AI, evaluation, and building alone, from a radiologist who reads every day and tests before he trusts. A few essays a month, nothing else.

Join for FreeFree · Unsubscribe anytime

Related

  1. Know the resolution of the instrument you trust

    A vision model attends in tiles of about three millimetres on a panoramic radiograph. Most of what a radiologist calls is smaller. The one-minute arithmetic.

  2. Three training runs, one lesson

    Three fine-tunings on my own radiology reads learned my report and not the image, and no training dial showed it. The three checks that catch it.

  3. Your own labels drift as you learn

    My facts per read nearly doubled over a reading campaign. A sparse answer key scores true findings as inventions. Measure your drift before you grade anyone.