Write the report so the patient's chatbot cannot inflate it
By Reza Motaghi

On this page
In 2018 a group in St Louis asked patients and radiologists the same question about the same phrases. How likely does each one make cancer sound. The patients put "probably" at the top.1 To them, "probably metastatic disease" was the most certain thing a report could say. The radiologists who write that phrase ranked it sixth. The radiologists' most certain phrase, "diagnostic for", the patients ranked third. The two groups agreed on exactly one thing. "Cannot exclude" sat at the bottom for both.
That last line matters more than it looks. The patient, reading alone, gets the hedge right. The report was written for one reader, the referring doctor, in a dialect that reader speaks. Then the patient learned to read it, imperfectly but not stupidly. And now there is a third reader, a machine that does not know which phrase is a routine hedge and which is a warning. It reads top-down and hedges nothing. So it inflates the hedge and headlines the incidental. The patient arrives already frightened, or already reassured, by a reading nobody in the chain wrote.
What "the three readers" means
A report keeps talking after you sign it. For a century it talked to one person, the doctor who ordered the study, and the whole craft of the report was built for that reader. The hedge that means unlikely. The "unremarkable" that means fine. The incidental finding tucked into the body because the referrer knows to skip it. That reader is now the minority.
The second reader is the patient, who by law in the United States sees the report the moment it is signed. The third reader is the patient's chatbot, which the patient pastes the report into within the hour. Three readers, one text. The three readers is a way of writing that does not choose between them. It says plainly what was found, what was not, and what each means. The machine cannot stretch a phrase and the two humans do not have to decode one. It costs a sentence. It saves the appointment.
The receipt: my own report, read by the third reader
I write these reports every day, and I test the models patients read them with. So I took one of my own reports, fully anonymised, and fed it to a frontier model with the question a patient would ask. What does this mean for me.
The hedge on a structure the image did not fully show came back as a possible diagnosis. An incidental finding, the kind the referrer skips, became the headline. Nothing the model said was invented. It did not hallucinate. It filled the context it did not have with the most likely reading, the way it fills any missing fact. To a machine that has never sat in a reading room, the most likely reading of "cannot be excluded" is that the thing might be there.
Then I rewrote the report and fed it again. The important negatives were stated. Routine was called routine in the sentence itself. The uncertain call stayed uncertain, because the image was, and the sure calls stopped hiding in hedges. The one thing that mattered came first, in plain words. The second version could not be inflated. There was nothing left to stretch. The tell is that the second version was also the better report for the referring doctor. Writing for the machine had made me write for the humans.
The patient reads first now
The order changed by law and then by habit. Since April 2021 American health systems have had to release results to the patient portal without delay.2 At one large system the share of imaging reports opened by the patient before the ordering doctor rose from 18.5 to 44 percent.3 The median wait from signing to the patient's first look fell from 45 hours to five and a half. Younger patients open faster.
They want it that way. In a survey of more than 8,000 patients, one in thirteen said immediate release had made them more worried.4 Among those with abnormal results it was one in six. Nearly 96 percent preferred immediate access anyway, and the abnormal-result group preferred it just as strongly. The reports they open are not written for them. Of 97,000 reports one group measured, about 4 percent read at or below an eighth-grade level, the reading level of the average American adult.5 So the patient does what anyone does with a text in a foreign register. They ask a machine. About a third of American consumers used an AI chatbot for health information last year, up from a sixth the year before.6 Nobody has yet measured how many paste a radiology report in. I would not bet against most.
What the third reader does to a hedge
The machine does not misread the report. It reads a report that was not written for it, and it does two measurable things.
It collapses uncertainty. Researchers built a benchmark of thousands of hedged clinical statements and asked language models to restate them. The original level of certainty survived between a third and a half of the time.7 About two in five hedges came back as definite statements. That is a preprint, one benchmark, and it says what my own test said.
And it over-calls. Three models read several hundred mammography reports and predicted malignancy from them. They caught more cancers than the radiologists and cleared far fewer of the benign cases.8 Their specificity ran between 29 and 43 percent. The radiologists' sat around 70. High sensitivity with low specificity is what a frightened reader looks like in numbers.
There is a second failure that runs the other way. Across 13,000 sentences of model summaries, the machine left out something that mattered about twice as often as it invented something.9 Omission is false reassurance. The report said "no fracture" and the summary did not, or the report hedged and the summary dropped the hedge in the reassuring direction. Both directions come from the same cause. The machine does not know which sentence carries the weight, so it guesses.
Now the honest counterweight, because it changes the rules below. Most summaries help. One department gave patients model-written summaries of their own reports. Ninety-two percent said the summary clarified the original.10 Their confidence that they understood it rose from a quarter to nearly all. A systematic review of sixteen studies found the summaries factually accurate most of the time.11 The harmful-error tail was a few percent once outliers were removed. Inflation is the tail, not the mode. But the tail is where the frightened phone call comes from, and the tail is written by the report.
The field has rewritten its language for a machine before
Radiology already did this once, on purpose. In the late 1980s the American College of Radiology began building a fixed vocabulary for mammography reports. The purpose was a database that could read every report the same way and audit the outcomes. It became BI-RADS, first published in 1993.12 The field rewrote its words for a machine reader and the humans got clearer reports as a side effect. That is the pattern I am asking for again, without the committee.
Because the committee has not finished. In 2025 the College's own quality commission wrote that the language of diagnostic certainty has never been organised into a commonly accepted framework.13 Radiologists, referrers and patients, it said, read the same phrases differently. Referrers and radiologists disagree with each other on what "consistent with" means, as radiology's own hedge studies keep finding.14 A thousand lay readers were asked what report sentences meant. Fewer than half of their answers were right, and a patient-summary format did better than the traditional one.15 The report was always a message to someone who was not in the room. Its language was never fixed. Now a machine reads it top-down, and the cost of the vagueness lands on the patient.
Five rules for the third reader
Five rules, each one a sentence in the report, and two of them changed after I read the evidence against my first version. I say where.
1. State the important negatives out loud. "No fracture. No mass." A chatbot cannot inflate what you have already ruled out, and it cannot omit a negative it has to summarise. Off-ramp: pertinent negatives, the ones the clinical question raised, not a list of everything absent. A report that rules out forty things buries the one that mattered.
2. Label routine as routine, in the sentence itself. "A common benign finding, no action needed." Not in your head, where the referrer would have supplied it. In the report, where the machine can read it. Off-ramp: routine is a clinical judgment, not a comfort word. If you would want the referrer to act on it, it is not routine. Calling it that to soothe the patient is the same error in the other direction.
3. Hedge only where the image is uncertain. "Cannot be excluded" for a real doubt. Where you are sure, say so. This is the rule I changed. My first version said drop the hedges. The hedging literature says why not. A hedge where the image is genuinely uncertain is the honest report, and the liability calculus around defensive hedging is contested, not settled.16 So the rule is the callable floor applied to the report. State what the image supports, hedge exactly where it does not, and never hedge to protect yourself. A defensive hedge becomes a diagnosis downstream. Off-ramp: if a hedge is real, keep it and add what would resolve it. "A follow-up scan in three months would settle this" is a hedge a machine cannot inflate.
4. Put the one thing that matters first, in plain words. The chatbot summarises top-down. So does the patient. When one group rewrote oncology reports in patient-friendly language, the reports got longer and the consultations got shorter, by about six minutes each.17 Off-ramp: plain first does not mean technical never. The referrer still needs the precise line, and it follows the plain one. Two sentences, two readers, one order.
5. Once, read your own report through a chatbot as the patient. Ten minutes, the way you would read anything cold, before you defend it. You will change how you write. This rule also changed, and it gained a fence. Fully anonymise the report first, or write a synthetic twin with the same shape and no patient in it. A real report with a real patient's details does not go into a cloud model, not to learn a writing lesson, not for anything. Off-ramp: once is enough. This is a calibration, not a workflow.
The two I changed are the third and the fifth. The first became "hedge where the image is uncertain" instead of "stop hedging". The second gained the anonymisation fence, because the test only works if the patient is not in it.

What the rules are not
They are not a patient-friendly translation stapled to the report. Some health systems now draft one at the portal, and that is a good lever. But it is a second document, and the patient's chatbot reads the first one. They are not a case against hedging. Real uncertainty belongs in the report, stated as uncertainty. And they are not the only lever. Release timing and portal design change what the patient meets, and a health system can pull those levers too. The report is the one lever the radiologist holds alone.
What I do
I write these reports every day, and I test the models patients read them with. The report was always a message to someone who was not in the room. There are three of them now, and one is a machine that reads top-down and hedges nothing. I write so that reader cannot get it wrong, and the other two get it right.
If you write things other people will paste into a chatbot, or build the models they paste them into, the research page says what I test. The contact page goes straight to my inbox.
References
Footnotes
-
Mityul MI, Gilcrease-Garcia B, Searleman A, Demertzis JL, Gunn AJ. Interpretive differences between patients and radiologists regarding the diagnostic confidence associated with commonly used phrases in the radiology report. American Journal of Roentgenology, 2018. Patients ranked "probably metastatic disease" as conveying the highest likelihood, radiologists ranked it sixth; radiologists ranked "diagnostic for" highest, patients third; "cannot exclude cancer" lowest for both. https://doi.org/10.2214/AJR.17.18448 ↩
-
Office of the National Coordinator for Health IT. Information blocking regulations, applicability date 5 April 2021. https://www.healthit.gov/faq/what-are-applicability-and-enforcement-dates-information-blocking-regulations/ ↩
-
Pollock JR, Petty SAB, Schmitz JJ, Varner J, Metcalfe AM, Tan N. Patient access of their radiology reports before and after implementation of 21st Century Cures Act information-blocking provisions at a large multicampus health system. American Journal of Roentgenology, 2024. 1,188,692 examinations, 388,921 patients: reports first accessed by the patient before the ordering provider 18.5 versus 44.0 percent; median time from report finalization to first patient access 45.0 versus 5.5 hours; 1.8 hours for patients under 60 versus 4.3 for 60 and over. https://doi.org/10.2214/AJR.23.30343 ↩
-
Steitz BD et al. Perspectives of patients about immediate access to test results through an online patient portal. JAMA Network Open, 2023. 8,139 respondents: 7.5 percent more worried, 16.5 percent among those with abnormal results, 95.7 percent preferred immediate access, 95.3 percent among the abnormal-result group. https://doi.org/10.1001/jamanetworkopen.2023.3572 ↩
-
Martin-Carreras T, Cook TS, Kahn CE. Readability of radiology reports: implications for patient-centered care. Clinical Imaging, 2019. 4,094 of 97,052 reports, 4.2 percent, at or below eighth-grade level. https://doi.org/10.1016/j.clinimag.2018.12.006 ↩
-
Rock Health, 2025 consumer adoption survey, about 8,000 US adults: 32 percent used an AI chatbot for health information, up from 16 percent in 2024. Ayre J et al. Use of ChatGPT to obtain health information in Australia, 2024. Medical Journal of Australia, 2025: 9.9 percent of adults in the previous six months. Neither is radiology-specific. https://doi.org/10.5694/mja2.52598 ↩
-
Du, Lu, Qu. Possible or definite? A benchmark for evaluating diagnostic uncertainty preservation in clinical text. 2026, preprint. 9,184 cue-target pairs from 1,200 documents; uncertainty retention 33.6 to 45.6 percent; 37.9 to 45.0 percent of retained hedges rewritten as definite. https://arxiv.org/abs/2606.18471 ↩
-
Dai et al. Performance of ChatGPT-4o, Claude 3 Opus, and DeepSeek-R1 in BI-RADS category 4 classification and malignancy prediction from mammography reports. JMIR Medical Informatics, 2025. 307 patients: model sensitivity 84.0 to 92.7 percent, specificity 28.9 to 43.0; radiologists sensitivity 68.0 to 80.7, specificity 68.3 to 76.1. https://doi.org/10.2196/80182 ↩
-
A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine, 2025. 12,999 clinician-annotated sentences: omission rate 3.45 percent, hallucination rate 1.47 percent. https://www.nature.com/articles/s41746-025-01670-7 ↩
-
Sunshine et al. Evaluating the quality and understandability of radiology report summaries generated by ChatGPT: survey study. JMIR Formative Research, 2025. 118 patients: 92 percent said the summary clarified the report, confidence in understanding rose from 26 to 98 percent, 83 percent of radiologist ratings said the summary represented the report very or extremely well. https://doi.org/10.2196/76097 ↩
-
Kwok, Franklin, Wastney. Simplifying radiology reports for patients using ChatGPT: a systematic review. Journal of Medical Imaging and Radiation Oncology, 2026. Sixteen studies: factual accuracy 78 to 100 percent, completeness 83 to 100, harmful or potentially harmful error rate 0 to 36 percent, 0 to 8 excluding outliers. https://doi.org/10.1111/1754-9485.70076 ↩
-
Burnside ES, Sickles EA, Bassett LW et al. The ACR BI-RADS experience: learning from history. Journal of the American College of Radiology, 2009. https://doi.org/10.1016/j.jacr.2009.07.023 ↩
-
Shinagare AB, Shankar PR, Chernyak V et al. Potential frameworks for communicating diagnostic certainty in radiology reports: from the ACR Commission on Quality and Safety. Journal of the American College of Radiology, 2025. https://doi.org/10.1016/j.jacr.2025.07.027 ↩
-
Gunn AJ, Tuttle MC, Flores EJ et al. Differing interpretations of report terminology between primary care physicians and radiologists. Journal of the American College of Radiology, 2016. https://doi.org/10.1016/j.jacr.2016.07.016 ↩
-
Cho JK, Zafar HM, Cook TS. Use of an online crowdsourcing platform to assess patient comprehension of radiology reports and colloquialisms. American Journal of Roentgenology, 2020. 47 percent of 812 interpretations by about 1,100 lay readers were correct. https://doi.org/10.2214/AJR.19.22202 ↩
-
Hiding in the hedges: tips to minimize your malpractice risks as a radiologist. American Journal of Roentgenology, 2019, https://doi.org/10.2214/AJR.19.21428. Marks, Mishra, Ormsby. Hedging in radiology: a discussion on the ethical and financial implications on the US health care system. Journal of the American College of Radiology, 2021, https://doi.org/10.1016/j.jacr.2021.02.030. Wallis A, McCoubrie P. The radiology report: are we getting the message across? Clinical Radiology, 2011, https://doi.org/10.1016/j.crad.2011.05.013 ↩
-
Yang et al. Enhancing physician-patient communication in oncology using GPT-4 through simplified radiology reports: multicenter quantitative study. Journal of Medical Internet Research, 2025. Word count 819 to 1,026; physician-patient communication time 1,116 to 745 seconds; comprehension score 5.5 to 7.8. https://doi.org/10.2196/63786 ↩
