Skip to content
Lab Notes3 min read

Your own labels drift as you learn

By Reza Motaghi

Your own labels drift while you learn. Grade a model against them and you will punish it for seeing what you had not yet learned to write.

I read every case in my training set myself. One board-certified oral and maxillofacial radiologist, every read taken to consensus with a second reader, over several weeks of reading to a written label contract. Then I turned the counting tool I use on the models onto my own reads: how many separate facts does each read assert. Earliest third, middle third, latest third.

The count nearly doubled from the earliest reads to the latest. Same eyes, same protocol, same scanner, same kind of case. What changed was me. Writing to a contract taught me to record things I had always seen and never written down: the state of a crest, the outline of a border, the absence that matters. My early reads are not wrong. They are written in an earlier dialect.

About a third of the held-out cases, the cases every model is graded on, were read in that early dialect.

That corrupts one axis of the grading in particular: fabrication. When a model writes a true finding that my early read never recorded, the judge scores an invention. A sparse answer key turns sight into a lie, and it does so quietly, in the direction that flatters the grader.

The fix is not to edit the old reads in place. It is a dated revision of the answer key, harmonized blind to every model's output, so the reference can be cited by version like anything else that changes. I built that revision as a merge that keeps every row, after a dry run showed that regenerating the reads from scratch would silently drop a third of them. Until the pass has run, every fabrication figure I quote carries this caveat.

In reader-study terms: the reference standard has a learning curve of its own, and it needs a version number.

Measure your own drift before you grade anyone. Anyone who labels long enough drifts, and the drift is not noise. It is learning, and it has a direction.

What I would do differently: measure the drift halfway through the reading campaign, not after it.

The check

Two checks before any label set becomes an answer key:

  1. Count the facts per label in your first tenth and your last tenth. If the last is far larger, your reference standard has a dialect problem, not a quality problem, and the early labels are undercounting what you saw.
  2. Version the reference standard before you grade anything. Revise it blind to every model's output, as a dated pass, and never edit an old label in place. A benchmark that cannot be cited by version cannot be trusted by anyone, including you.

About the author

Reza Motaghi is an oral and maxillofacial radiologist and Chief Innovation Officer who reads every day and builds, evaluates, and trains the imaging AI for it. He built CBCTScope, the first CBCT viewer with native AI-agent control, and writes about what a model did on a real case and what it did not.

The newsletter

Get the essays as they’re written.

Imaging AI, evaluation, and building alone, from a radiologist who reads every day and tests before he trusts. A few essays a month, nothing else.

Join for FreeFree · Unsubscribe anytime

Related

  1. Read a paper for the experiments it lets you skip

    One paper read twice struck a planned experiment and handed me a step. How to read a paper for the experiments it lets you skip.

  2. Know the resolution of the instrument you trust

    A vision model attends in tiles of about three millimetres on a panoramic radiograph. Most of what a radiologist calls is smaller. The one-minute arithmetic.

  3. Three training runs, one lesson

    Three fine-tunings on my own radiology reads learned my report and not the image, and no training dial showed it. The three checks that catch it.