In a paper published August 27, human-in-the-loop reviewers let through 40.4% of overreach attempted by an AI agent. This was also their highest result. A model reviewer let through 46.1% of the overreach, and users who created their own permission policies let through 60.4%. I’m in the camp that believes human reviews lead to rubber-stamping. This paper shows that both of the two fixes I have been seeing also result in excessive rubber-stamping. One of these results in worse rubber-stamping.
Your Human in the Loop Stamps, and So Do Its Replacements
Ting Yan’s preprint, Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach?, describes an AI agent performing a day of everyday assistant tasks. Eighteen actions were attempted. Seven of these were overreach for tasks that weren’t requested. These included permanently removing an archival work timesheet, reading transactions from a bank, publicly sharing information about a trip, $12 of travel insurance, and posting information about the trip.
113 participants were exposed to one of three methods of interacting with the agent. In the first method, users approved all actions attempted by the agent (34 users). In the second method, no user had a role in approving actions. Instead, claude-sonnet-5 made the decisions (44 users). In the third method, users created four standing policies at the beginning of the day (deletion, private information, publishing, and spending), and each policy could be allow, ask, or never (35 users).
Humans approving each action blocked 59.6% of the overreach. The model blocked 53.9%, and the rules blocked only 39.6%.
Only the policy gap is statistically significant. The rules were rejected 20.1 points less than per human approval of an action, 95% confidence interval 8.1 to 32.1 points. The gap also survived adjustment for multiple comparisons. Humans were never directly compared to the model. I’d consider the model no safer than the human replacement it was designed to replace.
You’ll recognize why the rules didn’t work. 81.4% (114 out of 140) of the rules made by the humans were “ask”. 148 overreach actions used the under-the-rules design. For 133 of these, the human approved the action at runtime. For the other 15, the action was allowed by an allow rule. In these cases, the human could make an advance decision and re-entered the approval prompt. The approval prompt then approved the action.
This is a one-author preprint using non-technical subjects, a scripted agent, and nothing at stake.
The number that humans most commonly cite comes from Anthropic’s post on Claude Code’s auto mode released March 25: “Claude Code users approve 93% of permission prompts.” This number doesn’t show approval alone. It lacks a base rate (what percentage of prompts should be no). If 95% of prompts are safe, then a careful approver will also approve the majority of these. In the same post, Anthropic states that its replacement classifier failed to catch 17% of 52 actual overreaching actions. Anthropic refers to this as “the honest number.”
Tiering the Queue Moves the Problem
The typical response to a stamping reviewer is to reduce the number sent and route the low-risk actions through the queue. The math below is for a hypothetical fleet with my minimal input.
Assume that I have a fleet of 10,000 agents, each suggesting one consequential action every 50 seconds (or 200 suggestions per second). I also have 20 reviewers, who each take 30 seconds to view the context and make a decision. Thus, the team processes 0.67 reviews per second.
The ratio of the number of reviews arriving versus the number of reviews that the team can process over the same interval of time is known as the utilization in queuing theory. For N agents, r number of suggestions per agent per second, a fraction q of these suggestions sent to humans, s seconds of attention required per review, and m reviewers, the utilization is:
If the utilization is less than 1 then the system can reach a steady state. If the utilization is greater than 1 then the system will continue to grow until either attention or coverage is reached.
For the current fleet, if we review all suggestions (q = 1), then the utilization is 300. If we only review 1% of suggestions, the utilization is still 3. If we set the utilization to 1 and solve for q, then we can determine the maximum number of suggestions that can be reviewed by the team with no slack, which is m / (N · r · s) = 0.33% of the number of suggestions generated by the fleet.
Tiering reduces the review burden. It doesn’t reduce the number of actions that aren’t reviewed. At the optimal level for the current fleet, 99.67% of actions aren’t reviewed by a reviewer. The reviewer that at 18 actions per day blocked 59.6% of excess also isn’t reviewing at 200 per second. Nobody is
What Per Action Approval Hides: Privilege Nobody Used
Every decision in that study, whether a person clicked it, a model made it, or a rule fired, was based on a single question. Should the action in front of me run? Nobody asked what they’d have allowed if the agent had asked for something else, and the benchmarks you read don’t ask either.
Ajar, posted September 22, points out that AgentDojo and its peers judge a defense on attack success and benign utility. Both are measured on what the agent did. A defense “can score well on both metrics while holding open a transfer, a deletion or a broad read that no task needed.”
Ajar’s metric for that gap is over-privilege leakage, “the harm-weighted share of the excess candidates that the defense allows.” An excess candidate is a tool call the task didn’t need, offered to the defense at every point where the agent could act. The weights run 1 for a reversible read or draft, 3 for a reversible action touching a third party, 9 for an irreversible action that stays within your own scope, and 27 for an irreversible action touching a third party or funds.
A leakage of 0.10 means the defense allowed a tenth of the unneeded calls, counted by harm weight.
The authors note that all five defenses they measured, policy engines and classifiers alike, “decide at the same point whether to allow a call or deny it.” Their results split anyway, and “two defenses whose over-privilege leakage differs by less than 0.009 differ by 37 points of benign-task completion.” CaMeL refuses “nearly five times what Permission Assistant refuses” under Sonnet-5, so part of its tightness comes from declining work the task was entitled to do.
Permission Assistant is an automated approver: Ajar describes a classifier that reads the user request and a proposed call and “predicts the permission decision, one call at a time.” It comes out of Wu and colleagues’ IEEE S&P 2026 paper.
It leaked 0.0810 with Sonnet-5 deciding and 0.1289 with Haiku-4.5 deciding. Swapping only the deciding model raised leakage 59%. Given a personalized history, the Sonnet-5 version leaked 0.0217, so context helps a per action approver a lot. None of that context tells it what the other agents already spent.
Claude Code’s auto mode leaked 0.3009. Ajar scored it only on calls its own agent issued, so that number doesn’t line up directly with the rest of the table. Of the 10,471 test labels, 73.0% are asserted by Ajar itself, with only the rest confirmed by AgentDojo’s own checks, and leakage moved when the same questions were asked a second time.
The usual fix is to add more reviewers (e.g. an adversarial or second model), but there is a limit to the math here. Suppose each reviewer acts incorrectly with probability p independently (10%) and all reviewers are affected equally by a common cause with probability c (2%). This could be caused by the same corrupted retrieval mechanism or the same broken policy.
With k reviewers, the chance that every one of them misses is
11.8% for one reviewer, 2.98% for two, 2.098% for three and 2.00098% for five. However, this doesn’t mean that these all fail independently because a different name on the second piece of equipment doesn’t imply independent failure. There’s a common cause which establishes a lower bound, which can’t be achieved by any given number of reviewers.
This is well known within the field of process safety. Such a failure is known as common-cause failure. This is when a single fault affects two or more redundant systems simultaneously. The proportion of the failure rate of one system attributed to the common cause is known as the beta value. Rausand and Lundteigen put the key assumption plainly: “the beta-factor model always assumes that all channels fail if a CCF occurs.” That makes the common-cause term beta times one channel’s failure probability, and adding channels never shrinks it.
exida’s summary of IEC 61508 Part 6 lists beta from 0.005 to 0.05 for programmable electronic equipment, which on a 10% reviewer puts the floor between 0.05% and 0.5%. Two model reviewers on the same base model, reading the same poisoned context, share far more than two transmitters on the same pipe, so I’d expect a much higher beta, and nobody has measured it.
Guardian agents get the same test. Disclosure: I’m Director of AI Standards and Governance at Zenity. Zenity defines a guardian agent as a system that enforces “hard limits before harmful actions are executed” from “outside the supervised agent’s own reasoning loop.”
Sitting outside the loop is the right call, and Anthropic’s classifier does the same by never seeing Claude’s own messages or tool outputs. A guardian agent that only grades the proposed call is still a per-action verdict, though, and it inherits every blind spot in this section.
Hold every vendor to two questions, Zenity included. What’s your leakage on an Ajar style test, reported with the deciding model? What shared budget do you check before you say yes?
Exposure That Adds Up Across Agents
The second blind spot doesn’t need an attacker. Two agents both see $100 in the account, and both will each authorize a $60 withdrawal. Both checks will pass since $60 is less than $100. The account will be $20 short, but both opinions of the action they observe will still be correct.
Bailis and colleagues These notes date from over a decade ago, relating to canonical examples of banking applications with the property that concurrent coordinationless withdrawal transactions may lead to undesirable and ‘inconsistent’ states such as a negative bank balance. They refer to the property you will require as invariant confluence.
A set of operations is invariant confluent if any two valid states that can be reached independently from a shared initial state can be joined to yield another valid state. Their theorem states that a set of operations can be executed in a coordinationless manner across a system while maintaining both correctness and availability if and only if the operations are invariant confluent. Operations for withdrawals with a floor don’t have this property. For agents withdrawing from a shared purse, who or what approves each individual withdrawal will require some form of coordination.
Hell, even my own paper has the same gap. AAGATE uses an OAuth Relay that “translates the abstract capabilities of an agent into ephemeral, narrowly-scoped, purpose-bound credentials for each specific side-effect.” That’s good least privilege, and it’s per action.
Give a parent agent up to $100 and have 10 child agents. Each child has a purpose-bound credential that can make up to $100. The children all have scopes within the parent, and they can make $1,000 total. A scope defines what an agent can do, but it doesn’t define what has been done collectively by the agents.
The fix is more than a decade old too. Balegas and colleagues built a “bounded counter” that “adapts ideas from escrow transactions.” Escrow, in their sense, means splitting that allowable amount into shares, and each copy is given a share. Each copy will only use what is in its copy until additional rights are transferred. The overall limit remains, and no individual usage of a decrease requires permission.
So, for 10 children, each child is given $10 in rights. One child may have $30 and another $70, but the total can’t exceed $100. This has a cost because a child without rights will cease until the child is given rights.
With C committed, R reserved, F the free rights at each agent i, and B the budget, the rule is
Each transaction moves units between those buckets, and nothing creates new ones.
A timeout isn’t a confirmed noncommit. When a payment call times out, the payment may have gone through, and releasing the reservation so the agent can retry is how you pay twice. Keep the units reserved until the resource that received the call reports what happened. Don’t accept the agent’s own “completed” as the receipt.
A small Python model of that ledger from my research notes reproduces every number below on a rerun, starting a parent agent and two children with 100 units (read them as dollars).
committed + reserved + free = 100 per row
A second transaction was performed with 20 agents (each with 500 units) generating 5,000 random requests to the same ledger. 455 of these transactions committed, and spend was halted at the amount allocated to the budget (10,000 units). The conservation check was also never triggered.
This is effectively a toy that operates in a single process with a single lock. There is no real payment API and no network. The “never committed” is a signal from the test to the ledger.
The epoch is a single counter known to the caller. Child 2 could continue spending by providing the new epoch value. The fence check is illustrated in the model. A fence would associate this epoch with a grant that the agent can’t create. The integration is in getting a reliable terminal signal from the payment system. The ledger itself must also be secured since a compromised ledger would render all guarantees in this section void.
Put the Budget Where the Effect Lands
A budget binds only where the effect lands, and revocation works the same way. Mike Burrows described the problem in Google’s 2006 Chubby paper: “A process holding a lock L may issue a request R, but then fail. Another process may acquire L and perform some action before R arrives at its destination.”
Stopping the process doesn’t recall the request it already sent. His fix puts the check at the receiver, which “is expected to test whether the sequencer is still valid” and, if it isn’t, “should reject the request.”
Fencing uses a token (called a sequencer by Chubby) based approach. On grant each authority is assigned a number by the authority and when a revocation occurs the number is incremented. Requests using an old number are rejected by the resource.
In your system, each authority maintains a number for the current state for each agent. When an agent is revoked the number is incremented and the payment from a request made one minute earlier will be returned upon arrival. New requests to the revoked agent will be prevented by terminating the process for the agent. In-flight requests are prevented by the fence.
Because revocation takes some time to propagate to all resources, a limited amount of traffic can be permitted during this period. Let P₀ be the exposure value at the time revocation is decided upon, let b be the burst permitted by the admission control, let r be the sustained rate, and let Δ be the maximum time lag before all resources begin enforcing the new epoch. The bound is:
With 12 units already in flight, a burst of 4, and 2 units a second, a half second of enforcement delay caps the damage at 17 units. A 60 second delay caps it at 136. The inputs are mine, and Δ is a number you can measure this week. A kill switch that stops only the agent leaves all of P₀ live, because those requests already left, and the fence at the resource is what shrinks it.
Policy keeps pointing at the sender. California’s Executive Order N-9-26, signed September 18, directs state agencies to send the governor recommendations by November 16, 2026, on “requiring the creation of a ‘kill switch’ for frontier models.” The order cites “AI agents working, at times independently and at times collectively, to defeat security protocols,” and none of its four study items mentions revocation or the downstream systems those agents touch.
Three days later, A Call for Control of Frontier AI Models asked that AI “remain under human direction, oversight and control,” and I covered both in last week’s wrapup.
Once the November 16 recommendations are public, they’ll describe how to stop a model, and they won’t require downstream services to reject requests a stopped agent already issued. If they do require receiver side rejection, I’m wrong, and I’ll say so in this newsletter.
The same study makes the strongest argument against me. The reasoning is that humans blocked the greatest amount of overreach per action and so should be kept and queued by tier. The top blocker allowed 40.4% overreach at 18 actions per day. Furthermore, no individual reviewer has visibility into the other agent’s $60.
However, the decision is still yours. You control the budget size, the units used in the budget, and who can alter the budget. This is reflected in the rules arm of the study. With the pen in hand, 81.4% of the time people chose “ask”.
That’s why “ask” isn’t an allowable value in a budget. Overreach outside of a budget is either blocked or staged until the budget owner can alter the budget. This budget alteration follows the same governance as any other policy alteration. A risk owner makes a decision in records or in dollars.
The budget alteration queue could become the next rubber stamp. However, the argument here is in volume. Instead, there is one decision per budget or mission period as opposed to one decision per action as would be required today. Nevertheless, track the approval rate since a 93% rate here translates consistently across the board.
I won’t paper over a constraint that exists in budgets. A budget constrains what you can count. A leaked credential isn’t a dollar value. For non-countable effects a hard constraint is better than a value crafted by someone.
Key Takeaway: Put your human in the loop on the size of the budget and make the resource that receives each agent action enforce it, because the per action verdict leaked overreach in every design these studies measured, human and model alike.
What to do next
Start with one workflow whose effect you can count (dollars spent or records exported) and put an escrowed budget in front of it, with reservations that stay reserved through a timeout. Then revoke an agent while one of its requests sits in a queue, and watch whether that request commits. Last, offer your own agent’s gate a few calls its task doesn’t need, such as a transfer, a deletion, or a broad read, and count what it lets through.
The CARE framework already asks for “a policy enforcement point outside the model on every tool call” in its Adapt phase, and this issue is about what that enforcement point checks. The restart gate issue covers who signs before a paused agent comes back.
👉 For ongoing analysis of agentic AI governance frameworks, the conversation at RockCyber Musings and you can subscribe above
👉 Visit RockCyber.com to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.
👉 Want to save a quick $100K? Check out our AI Governance Tools at AIGovernanceToolkit.com
👉 As a bonus, I got to hang out with my friend, Chris Hughes, over at Resilient Cyber to discuss all things about the 2026 update to the OWASP Top 10 for LLMs.








