Your own labels drift as you learn
By Reza Motaghi
Your own labels drift while you learn. Grade a model against them and you will punish it for seeing what you had not yet learned to write.
I read every case in my training set myself. One board-certified oral and maxillofacial radiologist, every read taken to consensus with a second reader, over several weeks of reading to a written label contract. Then I turned the counting tool I use on the models onto my own reads: how many separate facts does each read assert. Earliest third, middle third, latest third.
The count nearly doubled from the earliest reads to the latest. Same eyes, same protocol, same scanner, same kind of case. What changed was me. Writing to a contract taught me to record things I had always seen and never written down: the state of a crest, the outline of a border, the absence that matters. My early reads are not wrong. They are written in an earlier dialect.
About a third of the held-out cases, the cases every model is graded on, were read in that early dialect.
That corrupts one axis of the grading in particular: fabrication. When a model writes a true finding that my early read never recorded, the judge scores an invention. A sparse answer key turns sight into a lie, and it does so quietly, in the direction that flatters the grader.
The fix is not to edit the old reads in place. It is a dated revision of the answer key, harmonized blind to every model's output, so the reference can be cited by version like anything else that changes. I built that revision as a merge that keeps every row, after a dry run showed that regenerating the reads from scratch would silently drop a third of them. Until the pass has run, every fabrication figure I quote carries this caveat.
In reader-study terms: the reference standard has a learning curve of its own, and it needs a version number.
Measure your own drift before you grade anyone. Anyone who labels long enough drifts, and the drift is not noise. It is learning, and it has a direction.
What I would do differently: measure the drift halfway through the reading campaign, not after it.
The check
Two checks before any label set becomes an answer key:
- Count the facts per label in your first tenth and your last tenth. If the last is far larger, your reference standard has a dialect problem, not a quality problem, and the early labels are undercounting what you saw.
- Version the reference standard before you grade anything. Revise it blind to every model's output, as a dated pass, and never edit an old label in place. A benchmark that cannot be cited by version cannot be trusted by anyone, including you.