Skip to content
One-Person Startup10 min read

Why you must read the image before you open the AI answer

By Reza Motaghi

A seated figure reading a sheet of paper while the monitor beside them is covered by a cloth, drawn in a single white line on black, a minimalist illustration about reading cold before you look at the AI answer
On this page

European breast screening has run on two readers per mammogram for decades. The guideline that made it standard was specific about one thing.1 The second reader reads without knowing what the first one said. When a Dutch programme alternated the two ways month by month, the blinded second read caught more cancers, 83 percent against 76.2 Two pairs of eyes only count as two when the second pair has not seen through the first.

AI was sold to every reviewer as the second pair of eyes. It was installed as the first. The model's read is on the screen before you have looked at the image, the pull request, the clause, the draft. Once you have seen its answer you cannot unsee it. You stop reading the thing and start checking a claim about the thing. You find what you were told to find, and you stop.

What the cold read is

The cold read is the read you write down before you open the model's answer. Written, not held in your head. A finding you only thought is easy to revise once a confident answer arrives. A finding you wrote is a record you have to argue with. It is the oldest rule of double reading, applied to a new second reader: a second opinion is only worth something when it was formed alone.

Its five steps are below. The idea is one sentence, and it is why the steps come in that order. Whoever looks first does the reading. If the AI looks first, you are not a second reader. You are a signature.

What happens at the desk

I read images every day. For a while now I have also been running frontier vision models on my own cases, blind. The routine is fixed. I read, I write my findings, then I open the model's read. Then I go back to the image for every line where we disagree. The order is not a preference. It is what keeps me the reader.

I know what happens in the other order because I have felt it. With the model's answer open first, my eyes go where it pointed. If it names a finding, I confirm the finding. If it calls the study unremarkable, I read a study I already believe is unremarkable, and I read it faster. I am no longer looking for what is there. I am looking for whether it was right. Those are different tasks, and only one of them is a review.

The tell is the same one I described in the default patient. The failure does not look like a failure. It looks like agreement.

Where the pull comes from

The pull has been measured, and it is not subtle. In a Radiology study of AI-assisted mammography, a wrong AI suggestion pulled very experienced radiologists from 82 percent accuracy down to 46.3 Less experienced readers fell further, from 80 to 20.

Eye tracking shows the mechanism. When the AI had missed a cancer, readers looked at that cancer less often and for less time.4 And when the same cases were randomised to AI-first or human-first double reading, human-first came out ahead, 85 percent to 81.5

None of this is new to psychology. In the judge-advisor experiments of the 1990s, people who formed their own judgment before seeing advice ended up the most accurate.6 They were also the best calibrated. People who only chose after the advice narrowed their search around it. The advice was human then. The finding did not depend on that.

One small study tested the order directly with an AI tool. Veterinary radiologists read X-rays, some of them asked to write a first read before seeing the AI.7 Those who did agreed with the AI less, whether it was right or wrong, and rated it less useful. What they did not do was take longer. The first read cost no time. The cost objection to reading cold has been tested once, in one domain, and it did not hold.

Radiology already wrote the rule

The reading room worked this problem out before AI arrived, in two forms. The first is perceptual. Once a reader has found one abnormality, the search for the next one degrades. Radiology named this satisfaction of search in 1962.8 What you have already found shapes what you keep looking for. An answer handed to you before you look is a finding you did not even have to find.

The second form is procedural, and the two flagship AI screening programmes show it better than any lab study. In the German programme, the radiologist reads first, without the AI.9 The AI speaks only after the radiologist has called a case normal, as a safety net, and it tags the exams it rates as very unsuspicious. The programme ran across 463,094 women. Detection rose 17.6 percent with no rise in recalls.

The Swedish randomised trial chose the opposite order.10 The reader saw the AI's risk score for every exam, and its marks on the higher-risk ones, during the read. The AI also decided which exams needed two readers. Both programmes improved detection. They chose opposite orders.

That is the honest shape of the evidence. AI-first is not always worse. It is worse for one reviewer, alone, with an unvalidated model in front of them and no protocol around the two of them. Population screening with a validated tool inside a designed workflow is a different question. Someone decided, in advance, when the reader would see the AI, and measured what happened. That decision is exactly what is missing from the desk where a model's answer simply appears.

Beyond the reading room

The pull is not a radiology quirk. Developers with an AI coding assistant wrote less secure code than a control group, and believed it was more secure.11 Writers who drafted with an opinionated language model shifted their own stated views toward the model's position.12 Physicians given a frontier model as a diagnostic aid did not reason better than physicians without it.13 And a review of 106 experiments in Nature Human Behaviour found that human plus AI performed worse than the better of the two alone on decision tasks.14 The pairing only won on creation tasks.

One caution I owe the second reader of this article. Nobody has run the sequence experiment on code review, contract redlines, or edited drafts the way it was run on X-rays. What is established in those fields is the presence effect: the AI's output changes what the human accepts and how confident they feel. That the order of seeing it matters there too is an inference from the imaging and judge-advisor evidence. I state it as one.

The cold read, in five steps

The steps are a routine, not a philosophy. The pull is strongest when you are tired and the model sounds sure. Each step carries the evidence that shaped it, and two carry a limit.

1. Read cold, and write it down. Before you open the model, read the source and record what you found. The findings. The three things you expect the review to flag. The clauses that worry you. Written, so that a fluent answer cannot quietly revise your memory of what you thought. This is the step with the best evidence behind it. Independent judges are the most accurate and best calibrated, and the first read did not cost time. It has one off-ramp. It is for review and decision, not for creation. If the task is to draft from nothing, letting the model go first is where the human-plus-AI gains showed up. The cold read starts when there is something to judge.

2. Then open it, and mark every disagreement. Line by line against your written read. Not "does this look right", which is the question that anchors you. "Where does this differ from what I wrote." A model can be wrong by adding, wrong by omitting, and wrong by being right about a different case. Mark all three. One thing this step is not. Reading cold makes you agree with the model less, whether it is right or wrong. Disagreement is not accuracy yet. It is the list you take into step three.

3. Go back to the source, never to the confidence. Settle each disagreement against the image, the code, the clause. Never by which answer sounds more sure. The model was trained to sound sure and you were not. Two things raise the quality of this step. First, name what you are looking at. Readers who knew the first read came from a machine corrected it more often, 19 percent against 14.5 Second, a validated tool inside a designed protocol can add detection without becoming the reader. The German safety-net order is the model. It speaks after you have spoken.

4. Keep the disagreements. A list of where you and the model parted, kept over weeks, is the cheapest evaluation you will ever run. It shows where the model is blind, and where you are. It is also the only way to notice when a model update changes the behavior your review depends on without changing the tone of its answers. Mine is the raw material for the evaluation work I now do. Yours can be a text file.

5. No cold read, no review. If you opened the model first, you are not reviewing. You are signing. That is sometimes a fine thing to do, on a routine case with a validated tool, on purpose. It is not fine to do by accident and call it review. Decide, each time, whether your name belongs on it.

The cold read, five steps. 1, read cold. 2, then open it. 3, go back to the source. 4, keep the disagreements. 5, no cold read, no review.

Two things the steps are not. They are not an argument against AI reading first in every setting. The screening programmes above show a validated model, placed by design, improving detection on either side of the human. And they are not a call for more attention. Explanations and heatmaps alone do not reduce reliance on a wrong AI answer. Interventions that make the person think before the answer lands do.15 Sequence is one you can run on yourself, without a product change.

What I sign for

Radiology settled long ago what a signature means. You sign for what you saw, not for what you agreed with. AI did not change the rule. It only made the rule easy to break, because the answer now arrives before the looking does.

That is why my daily protocol is a fixed order and not a habit I hope to keep. It is also why, when I built my own CBCT viewer, the agent got the camera and never got the read. Both are the same decision. Whoever looks first does the reading, so arrange things so that it is you. If you build or evaluate models that people will review, and want to know how a daily reader tests them, the contact page reaches me directly.

References

Footnotes

  1. Perry N, Broeders M, de Wolf C, Törnberg S, Holland R, von Karsa L. European guidelines for quality assurance in breast cancer screening and diagnosis, fourth edition, summary document. Annals of Oncology, 2008. https://pubmed.ncbi.nlm.nih.gov/18024988/

  2. Klompenhouwer EG et al. Blinded double reading yields a higher programme sensitivity than non-blinded double reading at digital screening mammography. European Journal of Cancer, 2015. 87,487 screens, the two methods alternated monthly, programme sensitivity 83 against 76 percent. https://pubmed.ncbi.nlm.nih.gov/25573788/

  3. Dratsch T et al. Automation bias in mammography: the impact of artificial intelligence BI-RADS suggestions on reader performance. Radiology, 2023. Accuracy with a wrong suggestion 45.5 percent for very experienced readers against 82.3 with a correct one, and 19.8 against 79.7 for inexperienced readers. https://pubs.rsna.org/doi/10.1148/radiol.222176

  4. Taib AG et al. Automation bias in action: eye tracking of humans reading screening mammograms with and without AI prompts. Radiology, 2026. https://doi.org/10.1148/radiol.252590

  5. Sossavi et al. Human-AI interaction in a cancer-enriched double-reading breast screening cohort: diagnostic accuracy and second-reader behaviour. Cancer Imaging, 2026. Same cases randomised to AI-first or human-first double reading, accuracy 85.0 against 80.8 percent. Second readers corrected a wrong AI-initiated call more often when told it was AI, 19.1 against 13.6 percent. https://pmc.ncbi.nlm.nih.gov/articles/PMC12910764/ 2

  6. Sniezek JA, Buckley T. Cueing and cognitive conflict in judge-advisor decision making. Organizational Behavior and Human Decision Processes, 1995. https://doi.org/10.1006/obhd.1995.1040

  7. Fogliato R et al. Who goes first? Influences of human-AI workflow on decision making in clinical imaging. ACM Conference on Fairness, Accountability, and Transparency, 2022. Nineteen veterinary radiologists. Requiring a provisional read before the AI did not increase task time. https://arxiv.org/abs/2205.09696

  8. Tuddenham WJ. Visual search, image organization, and reader error in roentgen diagnosis. Radiology, 1962. https://doi.org/10.1148/78.5.694

  9. Eisemann N et al. Nationwide real-world implementation of AI for cancer detection in population-based mammography screening. Nature Medicine, 2025. The PRAIM programme, 463,094 women, detection up 17.6 percent, recall rate not increased. https://www.nature.com/articles/s41591-024-03408-6

  10. Lång K et al. Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI). Lancet Oncology, 2023. https://pubmed.ncbi.nlm.nih.gov/37541274/

  11. Perry N, Srivastava M, Kumar D, Boneh D. Do users write more insecure code with AI assistants? ACM Conference on Computer and Communications Security, 2023. https://arxiv.org/abs/2211.03622

  12. Jakesch M et al. Co-writing with opinionated language models affects users' views. ACM CHI Conference on Human Factors in Computing Systems, 2023. 1,506 participants. https://arxiv.org/abs/2302.00560

  13. Goh E et al. Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Network Open, 2024. Diagnostic reasoning score 76 percent with the model against 74 without, not a significant difference. https://doi.org/10.1001/jamanetworkopen.2024.40969

  14. Vaccaro M, Almaatouq A, Malone T. When combinations of humans and AI are useful: a systematic review and meta-analysis. Nature Human Behaviour, 2024. https://www.nature.com/articles/s41562-024-02024-1

  15. Buçinca Z, Malaya MB, Gajos KZ. To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. ACM Conference on Computer-Supported Cooperative Work, 2021. https://arxiv.org/abs/2102.09692

About the author

Reza Motaghi is an oral and maxillofacial radiologist and Chief Innovation Officer who reads every day and builds, evaluates, and trains the imaging AI for it. He built CBCTScope, the first CBCT viewer with native AI-agent control, and writes about what a model did on a real case and what it did not.

The newsletter

Get the essays as they’re written.

Imaging AI, evaluation, and building alone, from a radiologist who reads every day and tests before he trusts. A few essays a month, nothing else.

Join for FreeFree · Unsubscribe anytime