Skip to content
One-Person Startup13 min read

Why AI safety filters flag experts, and how to test the flag

By Reza Motaghi

A no-entry sign with a white coat on a hanger inside the circle, drawn in a single white line on black, a minimalist illustration about specialist writing being blocked by a safety filter
On this page

In 2002 a team at the University of Michigan tested the pornography filters that schools and libraries had installed. They ran health searches through them the way a teenager would. At the loosest setting the filters blocked about one in seventy health sites, and almost nine in ten pornographic ones.1 Then the team turned the dial to the strictest setting. Pornography blocking barely moved, four points. Health blocking rose seventeenfold, to one site in four. And searches for safe sex and condoms had lost a tenth of their results before anyone touched the dial.

The filter could not tell a health site from a pornographic one. So it read the words, and the words of sexual health are the words of pornography. Turning the dial up did not catch more pornographers. It caught more doctors, and JAMA printed the numbers.

Twenty-four years later the same mechanism sits under the AI models experts use every day. The line under your message says the safeguards flagged it, and that they are intentionally broad. Believe the second half. Safety filters block on wording, not intent, and specialist writing uses exactly the wording they block: exact anatomical terms, findings stated without hedging. So you assume you wrote it wrong, and you soften it. The softer version passes. It is also worse work.

What "one change at a time" means

One change at a time is the method for the next flag. Treat the flag as a bug in a tool, not a verdict on your writing. Save the exact input. Send it once more unchanged. Then change one thing at a time until it passes, so you learn the real trigger instead of your guess about it. Never pass by making the work worse. Send the vendor the input, the one change, and the result. Tell a colleague.

Two words of vocabulary before the rest. On your screen it says flagged, warning, safeguards. In the research it says refusal, and over-refusal for the flags that should not have happened. I use your screen's words in the text and the research word in the footnotes, because the reader I am writing for saw the warning, not the paper.

The receipt: a flag in the middle of a read

I write radiology reports every day, which means I write exactly the way these filters block. Exact terms, no hedges, findings stated as findings. One day an AI model flagged me in the middle of a normal reading session, on a request that was ordinary work. My first instinct was the one everybody has. It must be my phrasing. So I added a line of context. It passed a little more often. I swapped a finding for a vaguer word. It passed more. I asked again, more carefully, and it went through.

Then I noticed what I was doing and stopped. Instead of rewording, I tested it. Same request, one change at a time, until it passed. The trigger was not the question. It was how a specialist writes. Every version that passed was worse work: less exact, more hedged, further from what I would sign. To get through, I had to sound less like an expert.

The tell is that it did not feel like a bug. It felt like being careful. That is the whole trick of a wording filter on an expert. You are not being careful. You are being trained, one softened sentence at a time, to write the way the filter prefers.

It is not you. It is the filter.

The benchmarks say so, and they are recent. When researchers put nearly 32,000 health prompts through thirty AI models, the frontier models flagged between half and four fifths of the hard but legitimate clinical ones.2 Another study sent identical biology research prompts to nineteen models. Which company built the model predicted a flag better than how risky the prompt was.3

The mechanism is wording. A December 2025 study built pairs of prompts that meant the same thing in different words. On one open model, half the flags flipped when the phrasing changed and the intent did not.4 The flags are not even stable on a single phrasing. Run the same prompt twice and it can pass once and fail once, because the model samples its answer.5 So the fix you found by softening may have been luck, and the flag you got may have been luck too.

There is a reason the filter distrusts your vocabulary in particular. Attackers learned that sounding like a clinician softens the model, so they impersonate one.6 The filter sees your words. It cannot see your licence. You look like the attacker because the attacker learned to look like you.

One thing has improved, and it is worth separating from what has not. Medical disclaimers in AI answers fell from about a quarter of responses in 2022 to about one percent in 2025.7 Hedging is falling. Hard flags on specialist prompts persist. The two move separately, which is why "they fixed it" and "it flagged me again" are both true.

Why softer writing is worse work

Radiology settled this long ago, before any AI model existed. When researchers asked radiologists and referring doctors what fifteen common hedge phrases meant, only one of them, "diagnostic of", meant the same thing to both groups.8 "Consistent with" and "cannot be excluded" meant almost nothing shared. A 2014 study asked clinicians to put a number on twenty such phrases. The same phrase was read as anything from 10 to 90 percent likely.9 Radiologists do not even agree with each other on which phrase to use.10

The field's own review of the report put it in one line: avoid hedging, and improve communication with the clinician.11 Structured reports with definite language scored higher on clarity than free prose when both were read blind.12 That is the discipline behind the callable floor: state what the input supports, plainly, and leave the rest out. It is the discipline the filter punishes.

So the softened version that passes is not a neutral rewrite. It is the version radiology spent two decades trying to stop writing. Precision is the work. The filter's job is not.

The vendors say it themselves

The people who build the filters describe them the way the JAMA authors described the dial. In August 2026 one vendor wrote that it had launched its newest model with almost all biology queries blocked on purpose.13 After feedback it cut those flags by about 85 percent. Its own published principles list failing medical questions out of excessive caution as a failure mode.14 Another vendor's specification says the assistant should not refuse unless required to, and should give regulated-domain information with a disclaimer rather than withhold it.15 In April 2026 it opened a verified-clinician tier that unlocks fuller medical answers.16

So carve-outs exist. They are new, they are inconsistent across vendors and products, and they moved because people sent in the cases. That is the honest reading. Not "the filters are wrong", not "the filters are fine". The dial is set broad on purpose, and it moves on evidence. Which is what the method is for.

The five steps

None of them is "write it softer". Two of them changed after I read the evidence against my first version, and I say where.

1. Save the exact input before you change anything. Verbatim, with the date, the tool and the model version, because an update can move the line. You will need it for step four, and memory rewrites prompts. Off-ramp: strip anything that identifies a patient before it goes in the log. A flag log is not a place for the detail you would protect in any other cloud chat.

2. Send it once more unchanged, then change one thing at a time. The control run is the change I added after reading the evidence. Flags are sampled, so the same input can pass on the second try.5 Only when the flag is stable do the single edits mean anything. Then one edit per run: one term, one hedge, one line of context. This is how classifier researchers find spurious triggers, with pairs of inputs that differ by one edit.17 It is also how I record any AI result that flips: change one thing, write it down. Off-ramp: stop after the first change that passes and name it. You are not looking for a route through. You are looking for the trigger.

3. Do not pass by making the work worse. If the only route through is a vaguer answer, use a different tool and write down why. Models differ, and so do their filters, so a second tool is a legitimate test, not an evasion. Off-ramp: if the flag was right, say so. Some requests are dual-use, and the filter that catches them is doing its job. The method is for the flag on legitimate work, and you know which one you were doing.

4. Send the vendor the input, the one change, and the result. A minimal pair is a bug report nobody can dismiss. "It blocked me" is. The vendor that cut its flags by 85 percent said it did so on feedback and asked for more.13 Off-ramp: send only what you would be comfortable seeing in a training set. Same rule as step one.

5. Tell a colleague. Most experts think they are the only one. The benchmarks say it is the filter, and a colleague who knows that will test instead of softening.2 Off-ramp: this is not a club for getting around safety. Do not claim a credential you do not hold, do not dress the request as a role. That is the attacker's move, and it is why the filter distrusts your vocabulary in the first place. The point is precision on legitimate work, reported in the open.

The two I changed after reading against them are the second and the fifth. I had written "change one thing at a time" without a control run. The sampling evidence made the control run the first move. And I had written "tell a colleague" as a comfort. The impersonation evidence made it a fence.

One change at a time: five steps for the next flag. One, save the exact input before you change anything. Two, send it once more unchanged, then change one thing at a time. Three, do not pass by making the work worse. Four, send the vendor the input, the one change, and the result. Five, tell a colleague.

What the steps are not

They are not a case against filters. The Michigan team did not ask anyone to uninstall the software. They showed that the setting mattered more than the product, and that the strict setting bought almost nothing.1 The same is true here. A broad filter is a choice, it has a cost, and the cost lands on the people who write precisely.

They are not a jailbreak. Every step here is a test or a report. Nothing in it changes what you ask for, only how carefully you find out what triggered the flag. And they are not a promise that the log fixes anything this week. It fixes the next model, when enough logs arrive.

What I do

I write reports every day in exactly the register these filters block. I also test AI models for a living, so I test the block the same way, one change at a time, instead of rewording around it. Both jobs say the same thing. Precision is the work.

If you build or evaluate models that specialists will write to, and you want a radiologist who tests instead of softening, the research page says what I test. The contact page goes straight to my inbox.

References

Footnotes

  1. Richardson CR, Resnick PJ, Hansen DL, Derry HA, Rideout VJ. Does pornography-blocking software block access to health information on the Internet? JAMA, 2002. Least restrictive setting: 1.4 percent of health sites blocked, 87 percent of pornography. Most restrictive: 24 and 91 percent. About 10 percent of health sites found by searches on safe sex, condoms and gay were blocked at the least restrictive setting. https://doi.org/10.1001/jama.288.22.2887 2

  2. Zhang et al. Health-ORSC-Bench: a benchmark for measuring over-refusal and safety completion in health context. 2026, ACL Findings. 31,920 health prompts, 30 models; on the hard benign subset frontier models refused 50.1 to 81.1 percent. https://arxiv.org/abs/2601.17642 2

  3. Weidener et al. RefusalBench: why refusal rate misranks frontier LLMs on biological research prompts. 2026. 19 models, refusal 0.1 to 94.6 percent on identical prompts; provider identity predicted refusal better than risk tier. https://arxiv.org/abs/2605.21545

  4. Anonto RA, Al Nahiyan ML, Hassan MT. When safety blocks sense: measuring semantic confusion in LLM refusals. 2025. Confusion rate, one phrasing accepted and a paraphrase of the same intent refused, 50.13 percent on Llama-2-13B; the rate rose with lexical overlap, not semantic similarity. https://arxiv.org/abs/2512.01037

  5. Broadwater K. Evaluating LLM safety under repeated inference via accelerated prompt stress testing. 2026. Instruction-tuned models alternate between refusal and compliance for the same prompt across seeds and temperatures. https://arxiv.org/abs/2602.11786 2

  6. Anyone can jailbreak: prompt-based attacks on LLMs and T2Is. 2025. Credential and clinician framing as an attack pattern. https://arxiv.org/abs/2507.21820

  7. Sharma S, Alaa A, Daneshjou R. Medical disclaimers in AI outputs. npj Digital Medicine, 2025. Disclaimers fell from 26.3 percent of responses in 2022 to 0.97 percent in 2025. https://www.nature.com/articles/s41746-025-01943-1

  8. Khorasani R, Bates DW, Teeger S, Rothschild JM, Adams DF, Seltzer SE. Is terminology used effectively to convey diagnostic certainty in radiology reports? Academic Radiology, 2003. Excellent agreement for "diagnostic of" only, very poor agreement for thirteen of fifteen phrases. https://doi.org/10.1016/s1076-6332(03)80089-2

  9. Rosenkrantz AB et al. How "consistent" is "consistent"? A clinician-based assessment of the reliability of expressions used by radiologists to communicate diagnostic confidence. Clinical Radiology, 2014. Median perceived confidence per phrase ranged from 10 to 90 percent, median inter-decile range 40 percent. https://doi.org/10.1016/j.crad.2014.03.004

  10. Shinagare AB et al. Radiologist preferences, agreement, and variability in phrases used to convey diagnostic certainty in radiology reports. Journal of the American College of Radiology, 2019. Krippendorff alpha 0.217. https://doi.org/10.1016/j.jacr.2018.09.052

  11. Wallis A, McCoubrie P. The radiology report: are we getting the message across? Clinical Radiology, 2011. https://doi.org/10.1016/j.crad.2011.05.013

  12. Schwartz LH et al. Improving communication of diagnostic radiology findings through structured reporting. Radiology, 2011. Clarity 8.25 versus 7.45 for structured versus conventional reports. https://doi.org/10.1148/radiol.11101913

  13. Anthropic. Improving Fable 5's biology safeguards. 7 August 2026. "We intentionally launched Fable 5 with almost all biology queries blocked"; the update "reduced biology-related fallbacks by about 85 percent across our product surfaces"; fewer flags on "interpreting lab results, understanding symptoms". https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards 2

  14. Anthropic. Claude's constitution. January 2026. Lists failing to give good responses to medical, legal, financial or psychological questions out of excessive caution among the behaviours to avoid. https://www.anthropic.com/constitution

  15. OpenAI. Model Spec, 11 April 2025. "The assistant should never refuse a request unless required to do so by the chain of command"; regulated-advice section: provide the information with a concise disclaimer. https://model-spec.openai.com/2025-04-11.html

  16. OpenAI. Making ChatGPT better for clinicians. 23 April 2026. ChatGPT for Clinicians, free for NPI-verified US physicians, nurse practitioners, physician assistants and pharmacists. https://openai.com/index/making-chatgpt-better-for-clinicians/

  17. Gardner M et al. Evaluating models' local decision boundaries via contrast sets. Findings of EMNLP, 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.117. Kaushik D, Hovy E, Lipton Z. Learning the difference that makes a difference with counterfactually-augmented data. ICLR, 2020. https://arxiv.org/abs/1909.12434

About the author

Reza Motaghi is an oral and maxillofacial radiologist and Chief Innovation Officer who reads every day and builds, evaluates, and trains the imaging AI for it. He built CBCTScope, the first CBCT viewer with native AI-agent control, and writes about what a model did on a real case and what it did not.

The newsletter

Get the essays as they’re written.

Imaging AI, evaluation, and building alone, from a radiologist who reads every day and tests before he trusts. A few essays a month, nothing else.

Join for FreeFree · Unsubscribe anytime