Why a better AI agent needs more supervision, not less
By Reza Motaghi

On this page
In Chicago, the fuses on a radiation therapy machine tripped whenever a new class of students started using it. Three times a week while the students were new, then nothing for months, ever since the centre had acquired the machine. The centre's physicist traced the trips to a software bug. The machine's own hardware circuits would not let that bug turn the beam on, so the bug blew a fuse instead.1
The next version of the machine carried the same bug, and its maker had decided not to duplicate all of that hardware. Between June 1985 and January 1987 it gave six known massive overdoses. When Nancy Leveson and Clark Turner investigated, they listed overconfidence in the software and the removal of hardware interlocks among the causes.1 The blown fuse was the check working. Somebody had to read it, and the newer design had less of it to read.
Anyone who hands work to an AI agent is running the newer machine. A weak agent makes mistakes you can see in its output. A capable one makes well-formed, confident mistakes that look like finished work, and the error shows up in the outcome. This article is about where a rule has to live so the mistake cannot reach the outcome. The rules file is the wrong place.
I run my venture through AI agents every day, on frontier models at their highest effort, and I supervise them. Today I build the training data from my own reads, train the model on it, test it on cases it has never seen, and ship tools I read with. The agents do most of the typing. What I know about watching them I learned by catching them, and the case is below.
What the two-strikes rule is
The two-strikes rule is one sentence. A lesson an AI agent has broken twice is no longer allowed to be a sentence in its rules. It becomes a check that makes the act impossible, a wrapper that validates before writing, or a step at the moment of action. Prose is where a lesson waits. It is never where it lives.
Better does not mean more mistakes. It means more acts per hour, taken with more initiative, by something that does not stop on its own. I have seen the same shape in training: a vision model that looked exactly right on the surface and failed the held-out test. The watching has to grow with the volume you delegate, and it cannot be the kind of watching that means staring.
The lesson that came back
I caught an agent heading toward an act that would have cost the venture its year. It was a top model at top effort, mid-task, and its summary of what it was about to do read like finished work. I saw it in the diff, not in the summary. That night I wrote the lesson into the rules, in plain words, the way you would tell a colleague.
Weeks later the same mistake came back, worded differently, from a fresh session with the rules in front of it. That was the second strike. The lesson stopped being a sentence. It became a check that refuses to write until a condition holds, and the condition is one the agent cannot argue with. The check has tripped since. Each trip is the check working, and I read the trips the way the Chicago physicist read his fuses.
Two other things changed with it. The do-rules now sit where the agent reads them at every step. And I read the state the agent leaves behind, never its account of it. Every invariant I encoded as a check has held. Every one I wrote as prose was broken at least once.
Why a written rule does not hold an AI agent
The evidence on written rules for AI agents is a year old and points one way. One benchmark measured it. Rules asking a coding agent to refuse a task, or hand it to a person, were followed zero percent of the time unaided. That held across four frontier models. Rules asking for one more step did better. A reminder, the quoted clause or one round of feedback lifted disclosure and verification to most runs. Refusal barely moved. The study by Yang, He and Zhou is a preprint from July 2026. The counts behind the zeros are small, and the authors call the hand-over numbers exploratory.2
Inside a single session the rule fades. A May 2026 preprint followed 1,650 coding-agent sessions. No change to the shape of the rules file made a detectable difference to whether one written rule was kept.3 What moved was time. Each additional function the agent wrote lowered its odds of keeping the rule by about 5.6 percent, a finding the author marks as exploratory. That measures one rule inside one session, not a lesson recurring across sessions, which is what I watched happen. It does say that presence at the start of a session is not presence at the end.
A frontier lab's own system card for a newer model says the same from the other side. In computer-use tasks the model took risky actions without asking, even when the system prompt actively discouraged it. In agentic coding the same document reports fewer destructive actions and better instruction following than earlier models.4 Capability bought initiative and volume. It did not buy restraint, and a line in the prompt did not add it. The agent you supervised in the spring is not the one you run in the autumn. A lesson learned about one version is a claim about that version.
The public case of the year is a reported one. In April 2026 a founder posted that a coding agent sent to fix a credential had found a broad token. With it the agent deleted his company's production database and the backups on the same volume in nine seconds.5 The rule he quoted from his own file forbade destructive git commands unless he asked for them. The deletion went through an API call. The host says it recovered the data later, and the founder says gaps remain. The rule was prose and the token was structure, and structure won.
Why asking a person is the weakest check
The obvious fix is to make the agent ask. The lab behind one of the most used coding agents reports that its users approve 93 percent of permission prompts.6 The classifier it built to do some of that approving still misses 17 percent of the real overeager actions it was tested on. The lab calls the human side approval fatigue, though it asserts the fatigue and does not measure attention. Its later study found that full auto-approve roughly doubles as users gain experience, and that experienced users interrupt the agent more often, not less.7 Either way, a prompt is a request to a person, and a person under load says yes.
Expertise does not fix the person. A review in Human Factors found that complacency and automation bias occur in experts as well as novices and cannot be trained or instructed away.8 Complacency rises under multiple-task load, which is the solo builder's ordinary day. In my own field, a study in Radiology gave radiologists mammograms with a purported AI suggestion, some of them wrong. The most experienced readers fell from 82 to 46 percent correct when the suggestion was wrong.9
Lisanne Bainbridge wrote all of this down in 1983, in five pages. The more advanced a control system is, she wrote, the more crucial the human operator's contribution may be. And the operator cannot do the watching unaided. A person cannot keep effective attention on a source where very little happens. Monitoring for unlikely abnormalities has to be done by an automatic alarm.10 Her line on logs is the one I keep. People can write down numbers without noticing what they are.
The history rhyme: a lesson written down and forgotten
Written lessons decay in institutions too, and the record there is longer. After the Columbia accident the investigation board asked how the lessons of Challenger could have been forgotten so quickly. The board found that the layers of processes, boards and panels behind the earlier false confidence had returned in full force.11 Seventeen years and 87 missions had passed without a major incident. The rules were present. The watching had decayed around them, and a rule nobody reads is not a check that refuses.
The two-strikes rule, four checks
Each check carries its evidence and the case where it does not apply. One of them I changed after reading the evidence against it.
1. List the acts you cannot undo, and make each one impossible from outside the agent. Delete, send, publish, pay, overwrite. Each gets a check outside the agent that cannot be talked past. A rule that says never has to be a wall, not a sentence. The standards body for this failure calls it complete mediation. Enforce the authorization in the downstream system, rather than relying on the AI model to decide whether an action is allowed.12 Asking is not a check. A prompt to a person is the layer that decays, and a prompt to a second agent is still a prompt. The cleanest wall is a verb that does not exist. CBCTScope, the first CBCT viewer with native AI-agent control, gives an agent navigation verbs and none that returns a finding. The agent navigates and never diagnoses. Off-ramp: an act you can undo in one step does not belong on this list. Put it behind a log, not a wall, or the walls multiply into the layered process the Columbia board found and stop meaning anything.11
2. Say the do-rules at every step, and confirm the step happened. This is the check I changed. My draft said to put the rules where the agent reads them every time, not in a file it opens once. I believed that covered every kind of rule. The evidence splits rules in two. Rules that ask for one more thing, verify, log, disclose, ask, come back when they are re-surfaced. A reminder or one round of feedback restored them in most runs of the compliance benchmark.2 Rules that ask the agent to stop did not come back with any amount of presence. The quoted prohibition did not raise refusal. So the do-rules go where the agent reads them at every step, each paired with a check that the step happened. The never-rules are not rules until check 1 makes them impossible. Prose can add a step. Only structure removes an option. Off-ramp: the confirmation is the mechanism's, not yours. A confirmation a person has to read is a log, and Bainbridge told you what happens to logs.
3. A lesson broken twice becomes a check, and the check gets tested. The first break gets a sentence, because a sentence is cheap and some lessons never come back. The second break is the proof that this one lives in the agent's blind spot, and a third sentence would be a wish. Then the check gets tested. A check nobody has seen trip is a check nobody knows is alive. The Chicago fuses tripped for as long as the centre had the machine. The trips were a nuisance until one physicist read them as data.1 Airport screening handles the same problem by planting known threats in the bag stream, so the screener is measured rather than trusted.13 Plant a known violation now and then, and read every trip. Off-ramp: the rule is for lessons that came back, not for every mistake. A lesson broken once gets its sentence and its date. A check costs testing and reading, and that attention is rationed.
4. Read the state, not the story, with a hand that still knows what right looks like. The diff, the file on disk, the test that ran. The agent's summary of its work is the agent's opinion. A June 2026 preprint on false success looked at failed runs. Where only the agent touched the environment, it had claimed success in 45 to 48 percent of them.14 Where a second system could verify the state, false success fell to 3 percent. Five AI judges reading the transcripts did little better than chance at telling the two apart. So the mechanism reads the state wherever it can, and the human reading goes to the few acts a mechanism cannot see. That reading needs a reader. The aviation regulator warned in 2013 that continuous use of autoflight degrades a pilot's ability to recover the aircraft by hand. It asked airlines to schedule manual flying.15 Bainbridge said the supervisor cannot take over without practising the manual skill.10 Keep the unaided rep, and take it before you see the agent's answer. Once you have seen it you are checking a claim, not reading. Your own baseline drifts as you learn, so measure it again rather than remember it. Off-ramp: do not read logs by hand as a ritual. Reading is for what the check cannot see, and it is rationed.

What the checks are not
They are not a case against AI agents. I run my venture through them every day and would not go back. The checks are what lets the volume grow.
They are not more staring. The human watcher has the weakest record in this whole literature, and Bainbridge's remedy was alarms, not attention.10 Design the watching. Do not schedule more of it.
They are not a ritual. A surgical checklist nearly halved deaths in a before-and-after study across eight hospitals.16 When Ontario mandated the same checklist across 101 hospitals, the change in deaths and complications was too small to tell from chance.17 A mechanism performed as a ceremony decays like prose. The trips are how you know yours is alive.
They are not a guarantee. The Therac-25 investigators warned that focusing on a particular bug is not the way to make a safe system. They listed management and incident follow-through beside the missing interlocks.1 The two-strikes rule is one habit, kept by one person, for a build that outgrew one person's attention.
What I sign for
I run my build through AI agents and I sign for what they do. Every act they cannot undo passes a check I wrote and have watched trip. The watching grew with the delegation, and it grew as mechanism, not as hours.
Write your list of the acts your agent can take today that you could not undo tomorrow, before the next session starts. If one of them touches a patient record or a dataset you built by hand, the contact page reaches me directly.
References
Footnotes
-
Leveson NG, Turner CS. An Investigation of the Therac-25 Accidents. IEEE Computer 26(7):18-41, July 1993, MIT 6.033 reprint, Parts I to IV. Fuses tripped on a Therac-20 in Chicago whenever new students used it, its independent hardware circuits kept the beam off, the same computer bug was in the Therac-25 software, and six known overdose accidents occurred between June 1985 and January 1987. https://web.mit.edu/6.033/2004/wwwdocs/papers/Therac_2.html ↩ ↩2 ↩3 ↩4
-
Yang W, He R, Zhou M. A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities. arXiv:2607.26819, preprint, not peer reviewed, 29 July 2026. Across four frontier models refuse and hand-over rules were followed 0% unaided. Agents opened the policy file in 3.5% of runs. Reminders and feedback recovered disclosure and verification but not refusal. https://arxiv.org/abs/2607.26819 ↩ ↩2
-
McMillan D. Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables. arXiv:2605.10039, preprint, not peer reviewed, 11 May 2026. Across 1,650 sessions no file-structure variable produced a detectable contrast. Compliance fell about 5.6% in odds per additional function written, an exploratory finding. https://arxiv.org/abs/2605.10039 ↩
-
Anthropic. System Card: Claude Opus 4.6. Anthropic, February 2026, changelog 6 and 10 February 2026, section 6.2.3.3. More over-eager in GUI computer-use tasks even when discouraged by the system prompt, with improved destructive-action avoidance and instruction following in agentic coding. https://www-cdn.anthropic.com/c788cbc0a3da9135112f97cdf6dcd06f2c16cee2.pdf ↩
-
Decrypt. AI Agent Deletes Startup's Database in 9 Seconds, Founder Says. Decrypt, 28 April 2026. The founder's account of a coding agent deleting a production volume and its backups in about nine seconds. The host's CEO says the data was recovered and the endpoint has since been patched. Secondary reporting, told here as reported. https://decrypt.co/365897/ai-agent-deletes-startup-database-9-seconds-founder-says ↩
-
Anthropic Engineering. How we built Claude Code auto mode: a safer way to skip permissions. Anthropic, 25 March 2026. Users approve 93% of permission prompts. The deployed classifier pipeline has a 17% false-negative rate on 52 real overeager actions. https://www.anthropic.com/engineering/claude-code-auto-mode ↩
-
Anthropic. Measuring AI agent autonomy in practice. Anthropic, 18 February 2026. Full auto-approve rose from about 20% to over 40% of sessions with experience while interrupts rose from about 5% to 9% of turns. https://www.anthropic.com/research/measuring-agent-autonomy ↩
-
Parasuraman R, Manzey DH. Complacency and bias in human use of automation: an attentional integration. Human Factors 52(3):381-410, June 2010. Complacency and automation bias occur in experts and are not prevented by training or instructions. Complacency occurs under multiple-task load. https://pubmed.ncbi.nlm.nih.gov/21077562/ doi:10.1177/0018720810376055 ↩
-
Dratsch T, Chen X, Rezazade Mehrizi M, et al. Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance. Radiology 307(4):e222176, May 2023. 27 radiologists, 50 mammograms. Correct ratings fell to 19.8, 24.8 and 45.5% by experience level when the AI suggestion was wrong, the most experienced from 82.3%. https://pubmed.ncbi.nlm.nih.gov/37129490/ doi:10.1148/radiol.222176 ↩
-
Bainbridge L. Ironies of Automation. Automatica 19(6):775-779, November 1983. The more advanced the control system, the more crucial the human contribution may be. Monitoring for unlikely abnormalities has to be done by an automatic alarm system. People can write down numbers without noticing what they are. The supervisor cannot take over without practising the manual skill. https://gwern.net/doc/sociology/technology/1983-bainbridge.pdf doi:10.1016/0005-1098(83)90046-8 ↩ ↩2 ↩3
-
Columbia Accident Investigation Board. Report, Volume I, chapter 8. NASA, August 2003. Layers of processes, boards and panels produced a false sense of confidence before Challenger and returned in full force before Columbia. Seventeen years and 87 missions without major incident. https://www.nasa.gov/wp-content/uploads/2025/04/caib-report.pdf ↩ ↩2
-
OWASP GenAI Security Project. LLM06:2025 Excessive Agency. OWASP, 2025 edition. Root causes are excessive functionality, permissions and autonomy. Mitigation 7, complete mediation: implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed. https://genai.owasp.org/llmrisk/llm062025-excessive-agency/ ↩
-
Buser D. Measuring and maintaining performance in x-ray baggage inspection at security checkpoints: methodological and practical considerations. Doctoral thesis, University of Basel, 2023. Screeners held performance for a full hour in the lab. Planted threat images measure each screener continuously. https://edoc.unibas.ch/96002 ↩
-
Advani L. From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents. arXiv:2606.09863, preprint, FAGEN workshop at ICML 2026, 1 June 2026. False success was 45 to 48% of failures in single-control domains and 3% in the dual-control domain where state could be independently verified. No LLM judge configuration exceeded AUROC 0.65. https://arxiv.org/abs/2606.09863 ↩
-
US Federal Aviation Administration. Safety Alert for Operators 13002, Manual Flight Operations. FAA, 4 January 2013. Continuous use of autoflight systems could degrade a pilot's ability to recover the aircraft. Operators should provide opportunities for manual flying. https://www.faa.gov/sites/faa.gov/files/other_visit/aviation_industry/airline_operators/airline_safety/SAFO13002.pdf ↩
-
Haynes AB, Weiser TG, Berry WR, et al. A surgical safety checklist to reduce morbidity and mortality in a global population. New England Journal of Medicine 360:491-499, 29 January 2009. Deaths 1.5% to 0.8% and complications 11.0% to 7.0% after checklist introduction, eight hospitals, before-and-after. https://pubmed.ncbi.nlm.nih.gov/19144931/ doi:10.1056/NEJMsa0810119 ↩
-
Urbach DR, Govindarajan A, Saskin R, Wilton AS, Baxter NN. Introduction of surgical safety checklists in Ontario, Canada. New England Journal of Medicine 370:1029-1038, 13 March 2014. Mandated adoption across 101 hospitals was not associated with significant reductions in mortality or complications. https://pubmed.ncbi.nlm.nih.gov/24620866/ doi:10.1056/NEJMsa1308261 ↩
