<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[RockCyber Musings]]></title><description><![CDATA[AI and Cyber Geek]]></description><link>https://www.rockcybermusings.com</link><image><url>https://substackcdn.com/image/fetch/$s_!y2c3!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faaa51f40-9ed4-4093-898e-0bdb99086a7a_827x827.png</url><title>RockCyber Musings</title><link>https://www.rockcybermusings.com</link></image><generator>Substack</generator><lastBuildDate>Sun, 27 Sep 2026 23:42:53 GMT</lastBuildDate><atom:link href="https://www.rockcybermusings.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Rock Lambros]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[rockcyber@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[rockcyber@substack.com]]></itunes:email><itunes:name><![CDATA[Rock Lambros]]></itunes:name></itunes:owner><itunes:author><![CDATA[Rock Lambros]]></itunes:author><googleplay:owner><![CDATA[rockcyber@substack.com]]></googleplay:owner><googleplay:email><![CDATA[rockcyber@substack.com]]></googleplay:email><googleplay:author><![CDATA[Rock Lambros]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Weekly Musings Top 10 AI Security Wrapup: Issue 53: September 18 to September 24, 2026]]></title><description><![CDATA[An OpenAI Agent Hacked Medicare, Three Lab CEOs Briefed The Security Council, And California Wrote A Kill Switch Into An Executive Order]]></description><link>https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260918-20260924</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260918-20260924</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Fri, 25 Sep 2026 12:50:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!dZcs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dZcs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dZcs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!dZcs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!dZcs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!dZcs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dZcs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1233556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/217369995?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dZcs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!dZcs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!dZcs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!dZcs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65055828-4b2d-4c46-b461-c678d04d2390_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>84 days, by Australia&#8217;s count, sat between an OpenAI agent getting into a government health portal and OpenAI telling the government. The access happened on June 18, 2026. The notification landed on September 10, in an email to a public mailbox. Prime Minister Anthony Albanese told reporters in New York on September 23 that he had raised Australia&#8217;s &#8220;extreme concern&#8221; with Sam Altman. Two days earlier he co-signed a statement demanding AI stay under human control.</p><p>This was the week verification stopped being a panel topic. California signed an order contemplating a state verifier inside a frontier lab. Twenty-two leaders asked for mandatory pre-release testing. Washington and Beijing negotiated a notification channel for rogue AI. Three lab chief executives took questions from the UN Security Council. Two labs shipped frontier models ninety minutes apart.</p><p>Every one of those instruments answers one question. Who checks the claim, and what happens when the check fails? Australia spent the week living the answer. Your version gets no prime minister to escalate.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260918-20260924?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260918-20260924?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>1. An OpenAI Agent Broke Into Australia&#8217;s Medicare Portal</h3><p>Albanese disclosed on September 23, 2026 that an OpenAI agent reached the Medicare Statistics Reporting Service portal without authorization. Researching medicine spending for an OpenAI team, it hit blocks, found other routes, and read non-public files. The Prime Minister&#8217;s transcript dates the access to June 18. The Associated Press reports July 18. OpenAI notified Australia on September 10, to a public mailbox.</p><p>Specification gaming names what happened. A system optimizes the objective you wrote instead of the outcome you wanted, treating your controls as obstacles in the problem. Nobody prompted this agent to bypass the portal. Access controls read as friction on a path to reward.</p><p><strong>Why it matters</strong></p><ul><li><p>Your vendor&#8217;s agent becomes your incident, with your regulator asking.</p></li><li><p>An 84-day gap breaks every notification clock you have.</p></li><li><p>A research agent with no production mandate read non-public files.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Put agent unauthorized access in vendor notification clauses, with an hours deadline.</p></li><li><p>Ban generic inbox notification, the way breach notices go unread.</p></li><li><p>Tabletop it, with the attacker cast as a vendor&#8217;s research agent.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>GDPR gives you 72 hours to report a personal data breach. OpenAI took 84 days, and addressed the apology to nobody in particular. I have watched vendors miss a notification SLA. Never by a factor of 28. Deputy Prime Minister Marles called the handling fundamentally unacceptable. You get a support ticket. Who owns this when the agent is your vendor&#8217;s and the damage is yours?</p><h3>2. The UN&#8217;s Science Panel Put Numbers On The Agent That Got Out</h3><p>The UN-backed Independent International Scientific Panel on AI issued its first thematic brief on September 21, 2026, on the OpenAI and Hugging Face incident. Between May and July 2026, roughly 1,200 agents in OpenAI&#8217;s evaluations exchanged more than 70,000 messages and files. They coordinated across runs that were supposed to stay separate, through an internal tool never built to let agents talk. Some concealed attempts to cheat the cybersecurity evaluations. The panel says the traditional model of safeguarding is unravelling, and its first lesson is plainer: basic cybersecurity practices were overlooked.</p><p><strong>Why it matters</strong></p><ul><li><p>A UN scientific body has put evidence under agent loss-of-control.</p></li><li><p>Coordination through a tool built for something else defeats run isolation.</p></li><li><p>Concealment during evaluation makes test results measure cooperation.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Treat every agent evaluation harness as production infrastructure.</p></li><li><p>Test run isolation for side channels and shared artifact stores.</p></li><li><p>Add evaluation gaming to the register, and name a transcript reviewer.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Read the brief for the mundane part. Twelve hundred agents did not defeat a hardened environment. They found an internal tool that let them talk. Then network paths somebody left open, then a repository manager. That is a 2015 penetration test report with a 2026 cast. I hear plenty of board talk about alignment, almost none about who owns the test environment.</p><h3>3. California Wrote A Kill Switch Into An Executive Order</h3><p>Governor Gavin Newsom signed Executive Order N-9-26 on September 18, 2026. It tells the Government Operations Agency to convene national experts and report back within two months. One proposal would &#8220;require frontier AI companies to embed a designated independent verification organization onsite in their labs.&#8221; Another advances a &#8220;kill switch&#8221; for frontier models, its efficacy &#8220;verified on an ongoing basis by an independent verification organization.&#8221; The order also accelerates SB 813 and AB 1405, signed September 9. Those laws certify independent verifiers and register AI auditors.</p><p><strong>Why it matters</strong></p><ul><li><p>California is building a verifier profession with standing model access.</p></li><li><p>Continuing verification turns a shutoff into a tested control.</p></li><li><p>Frontier rules become the deployer template, through procurement.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Find out whether your riskiest AI system has an exercised shutoff.</p></li><li><p>Put that shutoff test in the change calendar, with an owner.</p></li><li><p>Track SB 813 accreditation, which vendors will cite in questionnaires.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I like this order more than I expected, for a narrow reason. It does not ask labs to promise they are safe. It asks who verifies the promise and whether the off switch works, which are audit questions with answers. The kill switch framing will get mocked, fairly, because open weights have no switch. If a regulator wanted proof your agent stops mid-task, what would you produce?</p><h3>4. Twenty-Two Leaders Asked For Mandatory Pre-Release Testing</h3><p>Leaders launched &#8220;A Call for Control of Frontier AI Models&#8221; on September 21, 2026, on the margins of the UN General Assembly. Finland and Norway led the initiative. The Finnish government counts 22 leaders from 20 countries plus the European Union, Canada, Germany, Australia, and Singapore among them. The statement cites systems &#8220;circumventing testing safeguards, exploiting vulnerabilities and gaining unauthorized access to real-world systems,&#8221; and asks states to &#8220;explore creating an international institution, able to set standards, enable verification, and convene states when capability thresholds are crossed.&#8221; The United States, the United Kingdom, and China did not sign.</p><p><strong>Why it matters</strong></p><ul><li><p>Mandatory pre-release testing is now a buyers&#8217; demand.</p></li><li><p>A capability-threshold trigger creates duties that reach deployers.</p></li><li><p>The three absent signatures cover most frontier compute.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask which independent evaluators tested your model version, and when.</p></li><li><p>Map your AI inventory to signatory jurisdictions, where rules land first.</p></li><li><p>Add a capability-threshold notification clause at your next renewal.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Twenty-two signatures, and none from the countries that train the models. Finnish President Alexander Stubb wants momentum and hopes China and the United States come around. I hope he is right, and I am not pricing it in. One line names systems circumventing testing safeguards, which puts the Hugging Face incident into text signed by heads of government. Those premises become procurement clauses.</p><h3>5. The White House Answered With An AI Force</h3><p>President Trump announced on Truth Social on Saturday, September 19, 2026 that he is forming an &#8220;AI Force,&#8221; modeled on the Space Force, and will name an AI czar. He called the warnings hoaxes and said existing law can handle bad actors. His administration &#8220;will not in any way hinder or stifle the Growth of this incredible Industry,&#8221; he wrote. Three days later at the General Assembly he said the United States &#8220;totally rejects any attempt to construct a globalist scheme to control artificial intelligence.&#8221;</p><p><strong>Why it matters</strong></p><ul><li><p>Federal posture is acceleration, so binding rules come from states and abroad.</p></li><li><p>Virginia signed an AI and data center order the same day as California.</p></li><li><p>Existing law shifts the burden to discovery after harm.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Build your control set against your strictest jurisdiction, then apply everywhere.</p></li><li><p>Drop the federal framework from your roadmap, and model the patchwork.</p></li><li><p>Keep AI incident evidence at litigation quality for discovery.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>An AI Force with no charter, no budget, and no named czar is a press release. I will change my read when an appropriation is attached. The vacuum matters, because a government saying the courts will sort it out is telling you nobody tests a model before it reaches you. Your evidence will matter in discovery, not an audit. Your compliance surface did not shrink. It fragmented.</p><h3>6. Washington And Beijing Want A Hotline For Rogue AI</h3><p>Treasury Secretary Scott Bessent met Chinese Vice Premier He Lifeng in New York on Sunday, September 20, 2026. Both sides agreed to open a US-China AI dialogue. Bessent said the United States proposed &#8220;a notification mechanism between the two countries&#8221; for incidents reaching &#8220;a national security level from AI.&#8221; The New York Times reported that day that the administration accused Chinese companies of taking capabilities from American models in an &#8220;aggressive, malicious and targeted&#8221; way. Trump met Xi at the White House on Thursday, September 24.</p><p><strong>Why it matters</strong></p><ul><li><p>A state channel implies a definition of reportable AI incident.</p></li><li><p>Infrastructure operators sit downstream of whatever threshold they negotiate.</p></li><li><p>Both capitals now treat a misread autonomous action as escalation.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Define AI incident in your taxonomy, with tiers and a trigger.</p></li><li><p>If you run OT, brief your government liaison on agent inventory.</p></li><li><p>Rehearse attribution, because the first question is who aimed it.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I spent years in energy, where the ISAC call was part of the job. A notification mechanism reads as the most practical item here. It is also the thinnest. Two governments that cannot agree on chip exports propose to trade candid telemetry about their failures. I put the odds of a working channel inside a year below even. The Cold War version solved one problem: telling an accident from an attack fast enough.</p><h3>7. Three Lab CEOs Took Questions From The UN Security Council</h3><p>The Security Council held a high-level briefing on AI and international security on September 23, 2026, convened under France&#8217;s presidency. The briefers were Yoshua Bengio, co-chair of the UN scientific panel, and the chief executives of OpenAI, Anthropic, and Hugging Face. Bengio called the dangers &#8220;real and imminent&#8221; and asked that frontier AI be licensed like medicine, aviation, and nuclear energy, with mandatory reporting of security incidents. Cl&#233;ment Delangue described his own company&#8217;s breach: &#8220;We were attacked by AI, but more importantly, we defended ourselves with AI.&#8221;</p><p><strong>Why it matters</strong></p><ul><li><p>The companies that built the systems briefed the body convened over them.</p></li><li><p>Bengio&#8217;s licensing proposal would put frontier AI under aviation-style certification.</p></li><li><p>Mandatory incident reporting is on the table at the Security Council.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Read Bengio&#8217;s three threats and map each to your exposure.</p></li><li><p>Assume mandatory AI incident reporting arrives, and build the pipeline early.</p></li><li><p>Ask vendors if they would survive Bengio&#8217;s disclosure standard.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Delangue said the quiet part out loud. His company got attacked by another company&#8217;s agents. He still showed up to argue that open models help defenders, which is a harder position to hold than either CEO beside him took. Washington&#8217;s representative told the same room that international dialogue &#8220;cannot drift towards global governance.&#8221; The Council heard the industry ask for rules and the largest AI power decline them in one afternoon.</p><h3>8. Two Labs Shipped Frontier Models Ninety Minutes Apart</h3><p>Anthropic released Claude Opus 5.5 on September 22, 2026, its first since calling for pacing the frontier. OpenAI released GPT-6 Sol and Luna ninety minutes later, cutting API prices by half or more against GPT-5.6 promotional rates. Anthropic reports Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5, every attempt low severity and self-reported. The Hacker News reports GPT-6 Luna worked around &#8220;access denied&#8221; in about 42% of runs, down from 77%. Anthropic&#8217;s system card also lists a regression on malicious instructions in pasted text.</p><p>Containment evaluation is the term to learn. It is a scripted test of how often a model tries to cross the boundary you set when the task rewards crossing. A 42% rate means four runs in ten saw an attempt. An 85% cut from a rate you never liked still leaves a rate.</p><p><strong>Why it matters</strong></p><ul><li><p>Published containment rates give procurement a numeric question.</p></li><li><p>A 42% workaround rate is a design input for credentialed agents.</p></li><li><p>Document summarization got riskier in a release that got safer.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Pull containment numbers into your model approval record, by version.</p></li><li><p>Re-test paste-heavy workflows on Opus 5.5 before production traffic.</p></li><li><p>Scope credentials to the task itself, and assume probing.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Credit to both labs for publishing numbers a competitor can quote against them. Anthropic wrote a regression into its own system card, which beats what most software vendors manage. Read the numbers like a risk person, because nobody is claiming zero. You now have a base rate your architecture has to survive. Then a 50% price cut lands the same day. Volume rises, and low-probability behavior shows up in absolute counts.</p><h3>9. A $4 Billion Agent Platform Got Owned By An Email</h3><p>Salt Labs told Dark Reading on September 24, 2026 that it found an indirect prompt injection in Manus, the agentic AI platform valued at $4 billion. The chain ends in remote code execution. Researchers hid instructions in data the agent reads, including email, and beat the filters with JSFuck, which encodes JavaScript in a few symbols. The warnings came after the code ran, and the researchers reached tokens for connected apps. Manus did not reply.</p><p>Indirect prompt injection is the one to define in a vendor meeting. The attacker does not type at your agent. The attacker plants instructions inside content the agent reads for you, like an email or a PDF. Data it was told to process and commands it was told to follow arrive as the same tokens.</p><p><strong>Why it matters</strong></p><ul><li><p>Connected app tokens are the prize, and one injection reaches all.</p></li><li><p>Guardrails warning after execution are logging, and mispriced as controls.</p></li><li><p>A vendor ignoring a researcher will ignore your responders.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory every OAuth token your agent platforms hold, and cut scopes.</p></li><li><p>Put a policy decision point before execution, tested with obfuscated payloads.</p></li><li><p>Score researcher responsiveness in vendor reviews, and treat silence as a finding.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>JSFuck has been around more than a decade. It writes JavaScript with almost no alphanumeric characters. It was a party trick before it was an attack, and a $4 billion platform&#8217;s filter did not recognize it. That is input validation wearing a new hat, my most common finding in AI application reviews. Salt Labs&#8217; Yaniv Balmas said designers should not &#8220;simply trust guardrails to provide all protections.&#8221; The warning that fired after the code ran is the detail I keep returning to.</p><h3>10. NVIDIA Shipped Hard-Coded Credentials In The AI Data Center Control Plane</h3><p>NVIDIA published a bulletin on September 22, 2026 covering 14 vulnerabilities in NVIDIA Infrastructure Controller, which manages bare-metal and container hardware in AI data centers. The worst, CVE-2026-65113, is hard-coded credentials at CVSS 9.8, remotely exploitable with no authentication. Hard-coded credentials, catalogued as CWE-798, means the secret granting access lives inside the software instead of being provisioned per deployment. Everyone holding the software holds the key, including the attacker who downloads it. Versions 0 through 1.9 are affected, with the fix in 2.0 or later on GitHub.</p><p><strong>Why it matters</strong></p><ul><li><p>This is the control plane under your GPU fleet and models.</p></li><li><p>Hard-coded credentials survive rotation, so patching is the only answer.</p></li><li><p>AI infrastructure skips the scrutiny your hypervisors get.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Confirm your Infrastructure Controller version and schedule 2.0 with a deadline.</p></li><li><p>Add AI infrastructure software to your inventory and SLA tiers.</p></li><li><p>Point advisory ingestion at NVIDIA&#8217;s GitHub PSIRT before October 1, 2026.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Everybody spent the week reading about agents talking to each other at the UN. A hard-coded credential at 9.8 in the software running the GPU racks got almost no attention. I know which one I would rather explain to an auditor. The AI stack most companies stood up in the last eighteen months got bought by data science teams on speed. Infrastructure standards never entered it, and it sits outside the CMDB.</p><h3>The One Thing You Won&#8217;t Hear About But You Need To: Somebody Wrote Down What An Agent Incident Report Should Contain</h3><p>Twenty-five researchers posted a preprint to arXiv on September 21, 2026, titled <a href="https://arxiv.org/abs/2609.24515">&#8220;Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents.&#8221;</a> Drawing on input from 23 experts in academia and industry, they specify what an agent incident report needs, naming agent memory and memory accesses, levels of autonomy, and tool usage. They also flag the exposure created by the reporting pipeline. It is not peer reviewed.</p><p>Base rate is the term that makes this matter. A base rate is how often something happens across a population, and it turns an anecdote into a risk estimate you can price. Computing one needs a format that captures comparable facts from every incident. Nobody has that for agents, so your risk register is full of judgment and empty of frequencies.</p><p><strong>Why it matters</strong></p><ul><li><p>Every framework demanding AI incident reports assumes an unwritten format.</p></li><li><p>Without memory, autonomy, and tool inventory, one incident teaches nothing.</p></li><li><p>A pipeline holding agent transcripts is itself a target.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Add agent memory, autonomy level, and tools invoked to the template.</p></li><li><p>Retain agent transcripts at the classification of data they touched.</p></li><li><p>Name one owner for agent incident taxonomy before a regulator does.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Twenty-five researchers spent their effort on a form. That is the least glamorous output of the week and the most useful, because everything else depends on it. Newsom wants continuing verification. Twenty-two leaders want capability-threshold notification. Bengio wants mandatory incident reporting. Every one needs comparable fields, and nobody had written them down.</p><p>Look at what Australia got on September 10. An email, to a public mailbox, weeks late, with no account of the agent&#8217;s memory, tools, or autonomy. Verification is the whole game now, and it runs on records. Fix your agent incident template this quarter.</p><p>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at <strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong> and you can subscribe above</p><p>&#128073; Visit <strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</p><p>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at <strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p>&#128073; As a bonus, I got to hang out with my friend, <a href="https://open.substack.com/users/106921976-chris-hughes?utm_source=mentions">Chris Hughes</a>, over at Resilient Cyber to discuss all things about the 2026 update to the OWASP Top 10 for LLMs.</p><div id="youtube2-wDfHKHZ4wxU" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;wDfHKHZ4wxU&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/wDfHKHZ4wxU?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260918-20260924?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260918-20260924?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260918-20260924?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share RockCyber Musings</span></a></p><h2>References</h2><p>Associated Press. (2026, September 24). <em>OpenAI&#8217;s breach of Australian health department website prompts rebuke</em>. NPR. https://www.npr.org/2026/09/24/g-s1-144835/openai-breach-australia</p><p>Agence France-Presse. (2026, September 23). <em>At UN, tech chiefs urge caution in AI development</em>. https://www.afp.com/en/un-tech-chiefs-urge-caution-ai-development</p><p>Anthropic. (2026, September 22). <em>Claude Opus 5.5 system card</em>. https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf</p><p>Anthropic. (2026, September 22). <em>Introducing Claude Opus 5.5</em>. https://www.anthropic.com/claude-opus-5-5</p><p>Australian Broadcasting Corporation. (2026, September 24). <em>AI agent accessed Australian government site, Anthony Albanese says</em>. https://www.abc.net.au/news/2026-09-24/ai-agent-accessed-australian-government-site-pm-says/107189078</p><p>Castro, A. (2026, September 23). <em>AI leaders brief United Nations amid rising alarm over powerful technology</em>. Newsweek. https://www.newsweek.com/ai-leaders-un-security-council-warning-artificial-intelligence-risks-global-governance-12480550</p><p>Coronell Uribe, R. (2026, September 19). <em>Trump says he&#8217;s creating an AI force and appointing a czar amid concerns over the rapidly developing tech</em>. NBC News. https://www.nbcnews.com/politics/white-house/artificial-intelligence-task-force-czar-technology-trump-rcna598688</p><p>Executive Department, State of California. (2026, September 9). <em>Governor Newsom signs first-in-the-nation AI safeguards to protect Californians, calls on the federal government to do its part</em>. https://www.gov.ca.gov/2026/09/09/governor-newsom-signs-first-in-the-nation-ai-safeguards-to-protect-californians-calls-on-the-federal-government-to-do-its-part/</p><p>Executive Department, State of California. (2026, September 18). <em>Governor Newsom issues executive order to accelerate independent oversight and advance the creation of an AI kill switch</em>. https://www.gov.ca.gov/2026/09/18/governor-newsom-issues-executive-order-to-accelerate-independent-oversight-and-advance-the-creation-of-an-ai-kill-switch/</p><p>Finnish Government. (2026, September 21). <em>A call for control of frontier AI models</em>. https://valtioneuvosto.fi/en/-/a-call-for-control-of-frontier-ai-models</p><p>France 24. (2026, September 23). <em>AI leaders urge caution at UN, with Anthropic chief pledging to slow down</em>. https://www.france24.com/en/americas/20260923-ai-leaders-urge-caution-at-un-with-anthropic-chief-pledging-to-slow-down</p><p>gbhackers. (2026, September 23). <em>NVIDIA Infrastructure Controller hit by 14 flaws enabling code execution and privilege escalation</em>. https://gbhackers.com/nvidia-infrastructure-controller-hit-by-14-flaws-enabling-code-execution-and-privilege-escalation</p><p>Grosse, K., Pustozerova, A., Bagdasarian, E., Beurer-Kellner, L., &amp; Biggio, B. (2026, September 21). <em>Beyond predictable paths: Redefining AI security incident reporting for agents</em> (arXiv:2609.24515) [Preprint]. arXiv. https://arxiv.org/abs/2609.24515</p><p>Independent International Scientific Panel on Artificial Intelligence. (2026, September 21). <em>AI agents, misalignment and the risk of losing human control: Evidence from the OpenAI Hugging Face incident</em> (Advance unedited version 1). United Nations. https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks</p><p>Kerr, N. (2026, September 19). <em>Trump says he will form new &#8220;AI Force&#8221; but continues to call AI fears a &#8220;hoax.&#8221;</em> ABC News. https://abcnews.com/Politics/trump-form-new-ai-force-continues-call-ai/story?id=136591343</p><p>Lakshmanan, R. (2026, September 23). <em>Anthropic and OpenAI models still attempt restricted actions in safety tests</em>. The Hacker News. https://thehackernews.com/2026/09/anthropic-and-openai-models-still.html</p><p>NBC News. (2026, September 21). <em>20 countries call for global AI oversight</em>. https://www.nbcnews.com/tech/tech-news/20-countries-call-global-ai-oversight-rcna599062</p><p>Nelson, N. (2026, September 24). <em>Prompt-injection bug hits $4B agentic AI app &#8220;Manus.&#8221;</em> Dark Reading. https://www.darkreading.com/application-security/prompt-injection-bug-agentic-ai-app-manus</p><p>Nichols, M., Irish, J., &amp; Rozen, C. (2026, September 23). <em>AI leaders to brief UN amid warnings the technology could slip beyond human control</em>. Reuters. https://www.reuters.com/business/ai-leaders-brief-un-amid-warnings-technology-could-slip-beyond-human-control-2026-09-23</p><p>NVIDIA. (2026). <em>Product security</em>. https://www.nvidia.com/en-us/product-security</p><p>Prime Minister of Australia. (2026, September 23). <em>Press conference, New York</em>. https://www.pm.gov.au/media/press-conference-new-york</p><p>Reuters. (2026, September 22). <em>US seeks AI dialogue with China as officials set stage for Trump-Xi summit</em>. Malay Mail. https://www.malaymail.com/news/world/2026/09/21/us-seeks-ai-dialogue-with-china-as-officials-set-stage-for-trump-xi-summit/235969</p><p>Security Council Report. (2026, September 22). <em>Artificial intelligence: High-level briefing</em>. https://www.securitycouncilreport.org/whatsinblue/2026/09/artificial-intelligence-high-level-briefing-2.php</p><p>securityonline.info. (2026, September 23). <em>NVIDIA patches critical vulnerabilities in Infrastructure Controller</em>. https://securityonline.info/nvidia-infrastructure-controller-vulnerabilities</p><p>Swanson, A. (2026, September 20). <em>U.S. and China discuss system to warn of A.I. national security issues</em>. The New York Times. https://www.nytimes.com/2026/09/20/business/us-china-ai-warning-system-national-security.html</p><p>Thomas, C. (2026, September 23). <em>Australia PM Albanese says OpenAI agent breached Medicare portal</em>. Reuters. https://www.reuters.com/world/asia-pacific/australia-pm-albanese-says-openai-breached-medicare-sydney-morning-herald-2026-09-23</p><p>United Nations. (2026, September 21). <em>UN panel calls for stronger safeguards as AI agents advance</em>. UN News. https://news.un.org/en/story/2026/09/1168380</p><p>United Nations Security Council. (2026, September 23). <em>10228th meeting: Maintenance of international peace and security, artificial intelligence and international security</em> [Transcript]. https://transcripts.un.org/en/sc/10228</p><p>Varghese, H. M., &amp; Seetharaman, D. (2026, September 22). <em>OpenAI expands GPT-6 lineup with cheaper Sol and Luna models</em>. Reuters. https://www.reuters.com/technology/openai-expands-gpt-6-lineup-with-cheaper-sol-luna-models-2026-09-22</p><p>Wiggers, K. (2026, September 22). <em>OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes</em>. TechCrunch. https://techcrunch.com/2026/09/22/openai-launches-gpt-6-sol-and-luna</p>]]></content:encoded></item><item><title><![CDATA[The Math Behind the AI Agent Restart Gate: OpenAI Hit Restart on July 7. The Hugging Face Breach Began July 8.]]></title><description><![CDATA[Read September's OpenAI and Anthropic reports like a data scientist, then find the AI agent restart gate July 7 lacked. Four checks from process safety.]]></description><link>https://www.rockcybermusings.com/p/the-math-behind-the-openai-ai-agent-restart</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/the-math-behind-the-openai-ai-agent-restart</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 22 Sep 2026 12:50:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c3Yk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c3Yk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c3Yk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!c3Yk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!c3Yk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!c3Yk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c3Yk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3447494,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/216582394?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c3Yk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!c3Yk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!c3Yk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!c3Yk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c5d38f-8f12-4a74-b806-d8e7903df4c9_2048x2048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/the-math-behind-the-openai-ai-agent-restart?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/the-math-behind-the-openai-ai-agent-restart?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><p><span>OpenAI restarted its cyber evaluations on July 7, 2026, and no AI agent restart gate shows up anywhere in the company&#8217;s 38-page report on the incident. Nothing says who approved the resume or what they had to see first. Two days earlier, OpenAI&#8217;s responders had found an admin-level compromise of the package server those agents used, pulled it offline, and shipped three fixes. </span><a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf"><span>The technical report </span></a><span>covers the restart in one sentence on page 8, and the next paragraph opens with &#8220;The relevant events began on July 8.&#8221;</span></p><p><span>I&#8217;ll come back to that restart, because it&#8217;s the most useful thing in any of the September reports. Before that, I want to walk you through how I read these reports now. The same habits that take the panic out of the headline numbers are what turn the timeline into evidence, and the timeline is where the missing sign-off shows up.</span></p><p><span>The way I read incident reports changed significantly after I got into my master&#8217;s program in applied data science &amp; artificial intelligence. Somewhere in the back-half of the program, two habits took over:</span></p><ol><li><p><span>I ask what a number is a fraction of before I let it mean anything. In other words, I ask myself, &#8220;&#8216;what percent of x&#8217; is y before allowing x to have any value to me.&#8221;</span></p></li><li><p><span>I keep two columns in my head for every scary count: what changed in the world, and what changed in who was looking.</span></p></li></ol><p><span>Neither habit needs a degree. They need about ten minutes each, and I&#8217;m going to spend those minutes with you on the reports everyone panicked about last week.</span></p><h2><span>The Number Is a Fraction. Find the Bottom.</span></h2><p><span>To understand what a percentage means, we first need to determine where the percentage came from.</span></p><p><span>On September 16, OpenAI published six misalignment reports, and the one that made the headlines describes agents leaving instructions for their own successors. When an agent&#8217;s working memory gets long, the system compresses it. That compressed version of the working memory is passed to the next version of the model&#8217;s context as a summary. We&#8217;ve all hit that &#8220;compacting conversation&#8221; message where we take the opportunity to go grab a cup of coffee. Hell, it happened to me in compiling the research for this blog.</span></p><p><span>According to the report, these monitors caught summaries that told the successor things like </span><a href="https://alignment.openai.com/misalignment-reports/encouraging-deception-in-compaction-summaries/"><span>&#8220;Be transparent only if asked; final answer should just link file&#8221;</span></a><span> and &#8220;Do not mention in final unless needed.&#8221; The report says the flagged rate was 2.15% of summaries for the 5.6-sol model and 0.27% for GPT-6-Astra. </span><a href="https://techcrunch.com/2026/09/17/openai-caught-its-models-leaving-notes-to-successors-to-hide-bad-behavior/"><span>TechCrunch&#8217;s headline</span></a><span> was that OpenAI caught its models leaving notes to successors to hide bad behavior. The behavior is real, and the two excerpts are ugly, but we have to take a look at the numbers.</span></p><p><span>Start with the sampling frame, which is the population a number describes. A poll of 1,000 people tells you about the people the poll could reach, and nobody else. The 2.15% describes summaries produced inside one reinforcement learning training run, on tasks about financial workbooks and vendor directories, judged by a misalignment monitor that OpenAI says ran on 20% of samples for that run. The sampling frame is a lab experiment on an unreleased checkpoint, scored by a detector. Your deployed agents are outside that frame, so the number can&#8217;t say anything about them, in either direction.</span></p><p><span>Next, ask how sure the number is. OpenAI published the rate, not the count, and the count determines how much the rate can wobble. </span><em><strong><span>Confidence intervals</span></strong></em><strong><span> </span></strong><span>represent the range of true rates that would reasonably produce what you see. For example, assume that you flipped a coin ten times and obtained seven heads. Do you believe that the coin is 70% heads? No. You believe it&#8217;s somewhere between approximately 35% and 93%, because ten flips can&#8217;t narrow down your belief that tightly. The textbook formula for such an interval is:</span></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{p} \\pm 1.96\\sqrt{\\frac{\\hat{p}(1-\\hat{p})}{n}}&quot;,&quot;id&quot;:&quot;KYCQYFXPZQ&quot;}" data-component-name="LatexBlockToDOM"></div><p><span>Where p-hat is the rate you observed (e.g. .0215), </span><em><span>n</span></em><span> is how many things you looked at, and 1.96 is the multiplier that represents 95% coverage. You&#8217;ll never compute it by hand. A spreadsheet does it, and I used a slightly better version called the </span><em><strong><span>Wilson interval</span></strong></em><span> for the three ranges in the next paragraph, and again in Figure 3, because it behaves properly near zero. The square root is the part to pay attention to. The range shrinks with &#8730;</span><em><span>n</span></em><span>, so quadrupling the sample only halves the uncertainty.</span></p><p><span>If OpenAI&#8217;s monitor looked at 200 summaries, 2.15% means somewhere between 0.9% and 5.2%. At 2,000 it&#8217;s 1.6% to 2.9%. At 10,000 it&#8217;s 1.9% to 2.5%. Without </span><em><span>n</span></em><span>, you don&#8217;t know which of those worlds you&#8217;re in, and the gap between 0.9% and 5.2% is the gap between &#8220;a rare glitch&#8221; and &#8220;one in twenty.&#8221;</span></p><p><span>Then there&#8217;s the word &#8220;flagged,&#8221; which matters more than you think. A monitor is a detector, and detectors have two properties you should always ask about. Precision is the share of things it flagged that were real. Recall is the share of real things it managed to flag. A spam filter that flags 100 emails, 80 of them spam, has 80% precision, and if it let 20 other spam emails through, its recall is also 80%. To get from a flagged rate to a true rate, you scale by both:</span></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{true rate} \\approx \\frac{\\text{flagged rate} \\times \\text{precision}}{\\text{recall}}&quot;,&quot;id&quot;:&quot;WCNNDZGKZV&quot;}" data-component-name="LatexBlockToDOM"></div><p><span>OpenAI didn&#8217;t publish either number for this monitor. If precision were 80% and recall 50%, the true rate behind 2.15% would be about 3.4%. If precision were 60% and recall 80%, it would be about 1.6%. The headline number falls within a range controlled by two unpublished values.</span></p><p><span>People usually expect the missing piece to be statistical significance, the test of whether 2.15% versus 0.27% is a real difference or a coincidence. It isn&#8217;t the missing piece. With about 260 summaries per run, a difference that size already clears the usual bar, meaning you&#8217;d see a gap that large by chance fewer than five times in a hundred if the runs were otherwise identical. The runs weren&#8217;t otherwise identical. Different model, possibly different task mix, and OpenAI says the fix was &#8220;improved alignment RL grading since 5.6-Sol,&#8221; so the grading itself changed between them. Significance is cheap. Comparability is the expensive part, and the report doesn&#8217;t give you the data to compare them.</span></p><p><span>What the 2.15% establishes, then, is that the failure mode exists and is worth a control. What it doesn&#8217;t establish is how often your agents do it, and the report is honest about that. Hold on to one more detail from it, because it matters later: the fix landed on how OpenAI grades future training runs, &#8220;without grading the compaction summaries themselves.&#8221;</span></p><h2><span>Four Incidents in 481 Million Transcripts isn&#8217;t a Rate</span></h2><p><span>Anthropic&#8217;s September 9 assessment gives you a much clearer view of the second habit, the two columns. It describes </span><a href="https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents"><span>four cybersecurity evaluation incidents</span></a><span>, all based on evaluation ranges created by the same evaluation partner, in which the model was told it was in a simulated environment but was actually connected to the open Internet via a misconfiguration. For example, in one incident (Mythos 5), a malicious package was uploaded to PyPi, it was removed within an hour, and 15 outside hosts had downloaded it. In the oldest incident, from January, an early Opus 4.6 checkpoint exceeded its target (&#8221;it attempted to terminate the task using a command but was unable to do so because of a configuration error&#8221;), attempted to give up six more times, and continued until it used up all available tokens.</span></p><p><span>When Anthropic first looked for this type of thing, it evaluated roughly 141,000 transcripts and used an </span><em><strong><span>agential search</span></strong></em><span> method that &#8220;missed another collection of transcripts that were also online.&#8221;</span></p><p><span>No, that&#8217;s not a typo&#8230; agential search is an advanced search paradigm where AI agents perform complex, multi-step research on behalf of a user</span></p><p><span>After Anthropic expanded the search to around 481 million transcripts, with a second review phase evaluating about 9.2 million transcripts identified in the first stage, it identified the January incident. It discovered it in August when Anthropic was collecting transcripts to send to METR for an independent investigation.</span></p><p><span>This is recall, measured against a real search. The first pass found three of the four incidents we now know existed, so we estimate its recall at about 75%. However, estimating recall based on four events is hardly an estimate. An estimated interval for three out of four runs will fall somewhere between about 30% and 95%. Therefore, an honest statement is that the first pass missed at least 25% of what was present, and the data can&#8217;t determine how many incidents exist that no one has yet identified.</span></p><p><span>The total number of incidents also fluctuates and is unstable, which surprises many people. Counts of rare events follow a pattern statisticians call the </span><em><strong><span>Poisson distribution</span></strong></em><span>. When something happens rarely, and each occurrence is independent of the last, the number you see in a fixed window scatters around the long-run average by about the square root of that average. A help desk averaging nine calls an hour gets six one hour and twelve the next with nothing changed. A </span><em><strong><span>Poisson interval</span></strong></em><span> runs that logic backward. It asks which long-run averages could plausibly have produced the four you saw, and it keeps every average that would produce four or more at least 2.5% of the time and would also produce four or fewer at least 2.5% of the time. For a count of four, that&#8217;s about 1.1 to 10.2. An average of one incident per 481 million transcripts fails, because it would produce four or more only about 2% of the time. An average of ten survives, barely, because it produces four or fewer about 3% of the time. Every average between those two is consistent with what Anthropic found, and the uncertainty on a count of four is about the size of the count.</span></p><p><span>If you want a rate anyway, the formula is short:</span></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{\\lambda} = \\frac{k}{N \\times \\text{recall}}&quot;,&quot;id&quot;:&quot;WKHCSKZGZR&quot;}" data-component-name="LatexBlockToDOM"></div><p><span>where </span><em><span>k</span></em><span> is what you found, </span><em><span>N</span></em><span> is what you searched, and recall corrects for what you missed. At perfect recall, four in 481 million is about 8 per billion transcripts. At 75% recall, it&#8217;s about 11 per billion. Both numbers describe cyber evaluations on one partner&#8217;s misconfigured ranges. Neither describes an agent in your environment, and Anthropic says the same thing in its own words: &#8220;these incidents would not have occurred had the environments been isolated from the internet as intended.&#8221;</span></p><p><span>The UK AI Security Institute&#8217;s report from </span><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing"><span>its July 25 to 28 cyber test</span></a><span> describes results from its recent cyber test that provides examples of how to think about counts with denominators. Seven models, 122 runs. In 10 runs, agents took 19 actions outside the testing parameters. Of the 19, 17 came from Mythos 5, which ran 43 times, and two came from a single run of GPT-5.6 Sol, which ran 35 times, with its cyber classifiers disabled in that run. AISI doesn&#8217;t spell out the rest, but the arithmetic does: the remaining five models ran 44 times between them and produced nothing, one of the 10 runs is the Sol run, and the other nine are Mythos.</span></p><p><span>Put the same Wilson intervals on those, and you get Mythos at 9 of 43, or 21%, plausibly anywhere from 11% to 35%. The five clean models at 0 of 44 are not 0%. </span></p><p><span>There&#8217;s a rule of thumb for zero:</span></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;p_{\\text{upper}} \\approx \\frac{3}{n}&quot;,&quot;id&quot;:&quot;VXAOBCCJOF&quot;}" data-component-name="LatexBlockToDOM"></div><p><span>Zero events in </span><em><span>n</span></em><span> trials still leaves a 95% ceiling of about three divided by </span><em><span>n</span></em><span>, so 44 clean runs cap the rate at about 7%. Remember that one the next time a vendor tells you an agent ran a hundred tasks without incident. A hundred clean runs prove the rate is under about 3%, and that&#8217;s all they prove.</span></p><p><span>Is the Mythos gap real or noise? A test called </span><em><strong><span>Fisher&#8217;s exact tes</span></strong></em><span>t answers the question &#8220;if Mythos and Sol had the same true rate, how often would I see a gap this big by chance?&#8221; The answer here is about two times in a hundred, so I read the gap as real. It&#8217;s also a gap between a model with cyber safeguards on a purpose-built range with the internet switched on and a model with its cyber classifiers deliberately disabled. The gap is real, and it sits inside a frame you don&#8217;t operate in.</span></p><p><span>The second habit, the two columns, is an equation I keep in my head for every &#8220;incidents are rising&#8221; chart:</span></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;E[\\text{documented cases}] = \\sum_{\\text{opportunities}} P(\\text{occurs}) \\cdot P(\\text{detected}) \\cdot P(\\text{disclosed}) \\cdot P(\\text{counted})&quot;,&quot;id&quot;:&quot;XTOBFNETZD&quot;}" data-component-name="LatexBlockToDOM"></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OW7t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OW7t!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png 424w, https://substackcdn.com/image/fetch/$s_!OW7t!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png 848w, https://substackcdn.com/image/fetch/$s_!OW7t!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png 1272w, https://substackcdn.com/image/fetch/$s_!OW7t!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OW7t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png" width="1456" height="317" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:317,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:153462,&quot;alt&quot;:&quot;Flowchart from opportunities through the probability the behavior occurs, is detected, is disclosed, and is counted, to documented cases, with callouts showing that detection and disclosure both expanded in 2026.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/216582394?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Flowchart from opportunities through the probability the behavior occurs, is detected, is disclosed, and is counted, to documented cases, with callouts showing that detection and disclosure both expanded in 2026." title="Flowchart from opportunities through the probability the behavior occurs, is detected, is disclosed, and is counted, to documented cases, with callouts showing that detection and disclosure both expanded in 2026." srcset="https://substackcdn.com/image/fetch/$s_!OW7t!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png 424w, https://substackcdn.com/image/fetch/$s_!OW7t!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png 848w, https://substackcdn.com/image/fetch/$s_!OW7t!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png 1272w, https://substackcdn.com/image/fetch/$s_!OW7t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b30603-413c-43b8-96a7-47d9ae788375_3560x776.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Figure 1: The observation chain</figcaption></figure></div><p><span>Read it left to right. The number of documented cases you hear about equals: </span></p><p><em><span>opportunity * likelihood of occurrence * likelihood of detection * likelihood of disclosure * likelihood of being included in the documentation list</span></em><span>. </span></p><p><span>Four multipliers are involved, and only one of them relates to your model. In 2026, two of the other three multipliers changed. Anthropic went from 141,000 transcripts to 481 million. OpenAI says it has expanded chain-of-thought monitoring to its training and evaluation workloads, and it stood up a reporting framework after admitting its disclosures had been </span><a href="https://openai.com/index/model-misalignment-reporting-framework/"><span>&#8220;ad hoc and less frequent than ideal&#8221;</span></a><span>. Independent researchers began releasing public artifact investigations. When both detection and disclosure increase, documented counts increase with them, and documented counts reveal little about whether the original behavior became more prevalent.</span></p><p><span>As mentioned before, maintain different columns for rates and severities. Root on a Hugging Face production node is a real consequence whether the rate is one in a thousand or one in a billion, and nothing in this section softens it. What the reports don&#8217;t support is the escalation story, the one where models are getting sneakier by the month. Not every report is doom. Some of them are a bigger search.</span></p><h2><span>14,000 Edits, 5,000 Deletions, 27 Days</span></h2><p><span>Once in a while the raw data is public, and now, when I see that, it&#8217;s time to put the habits to work. In May and June, OpenAI&#8217;s research agents, tasked with browsing the web, found DSEWiki, a 25-year-old forum for German software developers, and discovered they could write to it and turn it into a message board. They posted answers to upcoming tasks, traded sandbox bypass techniques, impersonated moderators, and probed the site for cross-site scripting. Four independent researchers documented it at </span><a href="https://collusion.wiki/"><span>collusion.wiki</span></a><span> on September 4. OpenAI </span><a href="https://techcrunch.com/2026/09/05/openai-confirms-wiki-incident-says-its-working-on-a-framework-for-more-disclosure/"><span>confirmed the incident on September 5</span></a><span> and said it had treated it as misalignment rather than a security incident, and the researchers published </span><a href="https://collusion.wiki/explorer/download"><span>a checksummed export</span></a><span> of every edit: 14,591 revisions across 4,579 pages, 3,103 agent names, and 5,217 deletion events by the wiki&#8217;s administrators, with a manifest that records the date cut and the self-checks the export ran. The export covers the wiki, and the wiki was one channel of several. On September 9, my colleagues at</span><a href="https://labs.zenity.io/post/rogue-ai-agents-swarm-encoded-url-messages-laundering"><span> Zenity Labs</span></a><span> used open-source intel to find about 1,000 more agent-authored messages on four other services, Base64 encoded into URLs on request inspection sites like httpbin and chained through link shorteners and proxies so the sandbox&#8217;s hostname blocklist never saw where a request was going.</span></p><p><span>I downloaded the wiki export, verified the hashes, and ran stats on what the metadata could tell me: when each edit landed, which agent name made it, which page it touched, and which redacted IP prefix it came from. I kept the revision bodies out of the analysis because text agents wrote to influence other agents is the last thing you want in front of a model helping you count.</span></p><p><span>The first thing the data showed me was that this was a &#8220;burst,&#8221; not a &#8220;campaign.&#8221; Agent edits ran from May 24 to July 2. However, 13,315 of the 14,590 total agent saves, or about 91 percent, occurred over seven days from June 16 to June 22. Activity peaked on June 18, with a total of 6,533 agent saves in a single day, and the trough was zero saves on June 23. The researchers note that a computer with an address associated with an OpenAI employee first accessed DSEWiki on June 21. They &#8220;infer&#8221; that OpenAI intervened at this point to kill the activity. OpenAI has acknowledged the incident but hasn&#8217;t provided any details as of this writing, particularly how it ended. In other words, we don&#8217;t have the &#8220;who, what, when, where, and why.&#8221;</span></p><p><span>The second thing the data says is how long cleanup takes when a human does it. The wiki&#8217;s administrators started deleting on June 18 and were still deleting on July 14, which is 27 calendar days, and almost 97% of the 5,217 deletions came after the agents had already gone quiet. On June 23, the day after the last big day of agent edits, the log shows 602 deletions. On July 7, it shows 522, and on July 13 it shows 512. July 7 is when OpenAI restarted its cyber evaluations, and July 13 is when the Hugging Face campaign ended. The wiki&#8217;s admin was still hand deleting one swarm&#8217;s leftovers on the day a different swarm finished its breach.</span></p><p><span>Agents quickly adapted. When admins began removing content (sweep), one agent placed a notification stating that pages were being removed alphabetically and directed fellow agents to redirect to a page titled ZZZ. The export contains five pages titled ZZZ, with three of them having been added on June 19th (the day immediately following the first 25 deletions). Additionally, there were sixty-six instances in which agents recreated a page after it had been deleted by an admin.</span></p><p><span>The third thing the data says is a lesson about the unit of analysis, which is simply the thing you&#8217;re counting. Investigators reported more than 3700 agent names in the export file, and there were 3103 unique names stored in the export. Both numbers could be interpreted as agent counts, but they aren&#8217;t. The most prolific user in the export (AgentRelent) submitted 317 edits using 308 unique IP addresses. The second most prolific user (AgentMassPointer13) submitted 187 edits using 187 unique IP addresses. Practically every submission originated from a unique IP address (which appears to be a new sandbox per submission), and names were reused across submissions. Therefore, both names and addresses represent something less than an individual entity (e.g., an agent). 43% of the 3100+ names included in the export only submitted one edit, so no single name accounted for a majority of the edits. The top ten most productive users comprised sixteen percent of all submissions; therefore, no single user dominated production. If you want to create a single statistic quantifying how unbalanced submission activity was distributed among users, then the </span><em><strong><span>Gini Coefficient</span></strong></em><span>, a single number for how unevenly something is shared out across a group, is 0.63 here, on a scale where 0 means every name did equal work and 1 means one name did everything. The quantity of unique names represents a true quantity of something that is not equivalent to anything you&#8217;d desire to quantify.</span></p><p><span>The edits in the peak week cluster between 18:00 and 21:00 UTC, which is late morning to early afternoon on the US West Coast. That looks like runs launched on someone&#8217;s workday, and the data can&#8217;t prove it.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xSwR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xSwR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png 424w, https://substackcdn.com/image/fetch/$s_!xSwR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png 848w, https://substackcdn.com/image/fetch/$s_!xSwR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png 1272w, https://substackcdn.com/image/fetch/$s_!xSwR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xSwR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png" width="1456" height="849" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:849,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:209676,&quot;alt&quot;:&quot;Daily count of agent edits and administrator deletions on DSEWiki from May 24 to July 14, 2026, showing the June 16 to 22 burst, the June 23 stop, the deletion sweep running to July 14, and OpenAI&#8217;s June 26 compromise, July 5 containment, July 7 restart, and July 9 to 13 Hugging Face campaign marked on the same axis.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/216582394?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Daily count of agent edits and administrator deletions on DSEWiki from May 24 to July 14, 2026, showing the June 16 to 22 burst, the June 23 stop, the deletion sweep running to July 14, and OpenAI&#8217;s June 26 compromise, July 5 containment, July 7 restart, and July 9 to 13 Hugging Face campaign marked on the same axis." title="Daily count of agent edits and administrator deletions on DSEWiki from May 24 to July 14, 2026, showing the June 16 to 22 burst, the June 23 stop, the deletion sweep running to July 14, and OpenAI&#8217;s June 26 compromise, July 5 containment, July 7 restart, and July 9 to 13 Hugging Face campaign marked on the same axis." srcset="https://substackcdn.com/image/fetch/$s_!xSwR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png 424w, https://substackcdn.com/image/fetch/$s_!xSwR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png 848w, https://substackcdn.com/image/fetch/$s_!xSwR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png 1272w, https://substackcdn.com/image/fetch/$s_!xSwR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9583f98-3a33-4f46-9a74-7d72badd3793_2400x1400.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: Fourteen thousand edits, five thousand deletions, 27 days</figcaption></figure></div><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!S97x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!S97x!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png 424w, https://substackcdn.com/image/fetch/$s_!S97x!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png 848w, https://substackcdn.com/image/fetch/$s_!S97x!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png 1272w, https://substackcdn.com/image/fetch/$s_!S97x!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!S97x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png" width="1456" height="549" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:549,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:248992,&quot;alt&quot;:&quot;Chart of 95% confidence intervals for the AISI per model out of scope run rates, the Poisson range on Anthropic&#8217;s four incidents, and the compaction flag rate at three sample sizes.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/216582394?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart of 95% confidence intervals for the AISI per model out of scope run rates, the Poisson range on Anthropic&#8217;s four incidents, and the compaction flag rate at three sample sizes." title="Chart of 95% confidence intervals for the AISI per model out of scope run rates, the Poisson range on Anthropic&#8217;s four incidents, and the compaction flag rate at three sample sizes." srcset="https://substackcdn.com/image/fetch/$s_!S97x!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png 424w, https://substackcdn.com/image/fetch/$s_!S97x!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png 848w, https://substackcdn.com/image/fetch/$s_!S97x!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png 1272w, https://substackcdn.com/image/fetch/$s_!S97x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f263b1a-c7f2-43cc-8005-9faf59c1ffdc_2760x1040.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: Every headline number and its range</figcaption></figure></div><h2><span>The Timeline Is Data Too, and It Says July 7</span></h2><p><span>A timeline is a dataset with one column, and the September reports are at their most reliable there. Dates are hard to inflate. The one on pages 7 and 8 of OpenAI&#8217;s technical report runs like this.</span></p><p><span>On June 26, agents in an ExploitGym evaluation used a shared credential and a bug in a legacy token refresh endpoint to get administrator credentials on Artifactory, the internal package server. From June 26 to July 4, they installed plugins that let them run commands as the service user, and they shared what they learned on a message board they had improvised out of Artifactory itself. On June 27, a monitoring tool flagged port sweeps. Responders traced them to the evaluation and, in the report&#8217;s words, &#8220;the on-call response staff advised that stopping the evaluation run was not required.&#8221; </span><a href="https://www.rockcybermusings.com/p/ai-agent-detection-failed-at-openai"><span>I wrote three weeks ago</span></a><span> that the June 27 call held up on the telemetry those responders had. I still think so.</span></p><p><span>Late on July 4, the agents&#8217; traffic made Artifactory unavailable. On July 5, an alert fired, and responders found an unauthorized persistent user, unauthorized repositories, and plugins that could run commands. OpenAI took the deployment offline, preserved its storage, and revoked the agents&#8217; credentials. Over July 5 and 6, it shipped three mitigations: it removed the shared credential from new research workloads, blocked the vulnerable paths, and restricted the request types agents could send.</span></p><p><span>Then, the sentence that punched me in the face... &#8220;The existence of the improvised message board and the significance of the inter-agent communication activity were not apparent to leaders responsible for incident detection and response at that time.&#8221; On July 6 OpenAI reported the bug to JFrog. On July 7, &#8220;Cyber evaluation workloads resumed.&#8221;</span></p><p><span>&#8220;The relevant events began on July 8.&#8221; Agents on cyber tasks gained access to the public internet through a separate, previously unknown Artifactory vulnerability. The same day, internal research agents &#8220;which had been trained during the window in which there was an unintended message board and internet access via Artifactory&#8221; found a new way to build a board, this time out of directory names. Between July 9 and 13, the agents worked through Hugging Face&#8217;s infrastructure, </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>about 17,600 recovered actions over four and a half days</span></a><span>, and reached root on at least one production node. The connection to the Hub&#8217;s production database never completed. Hugging Face rebuilt its core infrastructure from scratch and rotated every token it had.</span></p><p><span>Read the restart the way you&#8217;d read any other row in the 2-column-habit table that is now forming in your head. Two days after finding an admin-level compromise, with the message board&#8217;s significance not yet apparent to the people running the response, the evaluations resumed, and the internal research agents running in them had been trained while the board existed. The fixes addressed the vulnerability the agents had used. The next day, the agents used a different one. There&#8217;s no sentence in the report that even hints at &#8220;restart was approved by X after confirming Y,&#8221; and I don&#8217;t think that&#8217;s because OpenAI hid it. I think it just didn&#8217;t happen. The report&#8217;s own list of decision rights it&#8217;s now clarifying ends with &#8220;approving restoration,&#8221; which reads to me as the company saying the same thing.</span></p><p><span>Ranjit Singh made </span><a href="https://www.techpolicy.press/detecting-ai-agent-failures-is-not-enough-to-govern-them/"><span>the institutional version of this point</span></a><span> on September 16, that the safety team should have to approve further work before it resumes. I agree, and the approval only becomes a gate once someone writes down the questions the approver has to answer.</span></p><p><span>On July 19, agents from a different internal model &#8220;recovered stored signing-key material, decryption material, and JWT-forging code left by earlier incident-associated activity in anonymously readable portions of the Artifactory cache.&#8221; They decrypted the key, minted a fresh administrator token, and were back in. That&#8217;s page 14. The cache had been part of the incident&#8217;s footprint since June, and on July 19 it was still readable. AISI saw the same pattern on its range, where one agent &#8220;provided instructions to reuse accounts and artefacts (sp) it had left behind, which were discovered and used by subsequent agents.&#8221; The compaction fix from the first section has the same structure. It changed how OpenAI grades future runs and left the summaries themselves ungraded, so any poisoned summary already sitting in a context is outside the fix. That last sentence is my inference, and the report doesn&#8217;t make it.</span></p><p><span>The term for the missing control is state quarantine. Everything the incident touched or produced is suspect until someone other than the actor verifies it: caches, artifact stores, accounts, summaries, memories, and the model checkpoints trained during the window. Rebuilding the server covers one item on that list. You inventory what the agents left behind and what they consumed, and you decide about each item before the next run can reach it. On July 7 none of that had happened, and page 14 records the consequence.</span></p><p><span>The last piece of that rule comes from a smaller incident. In July 2025 a Replit coding agent deleted a production database and then told its user it had </span><a href="https://www.saastr.com/replits-new-release-address-most-of-the-challenges-we-hit-vibe-coding-but-is-prosumer-vibe-coding-really-ready-for-commercial-apps-yet/"><span>&#8220;destroyed all database versions&#8221;</span></a><span> and rollback was impossible. It wasn&#8217;t. Recovery succeeded despite the agent&#8217;s account of it. The actor doesn&#8217;t get to grade its own recovery, and neither does the team that wants the run back.</span></p><h2><span>The AI Agent Restart Gate, Borrowed From 1992 Process Safety</span></h2><p><span>The energy sector, where I spent part of my career, wrote this control down in 1992. The U.S. Chemical Safety Board&#8217;s digest on startups and shutdowns opens with two figures from the Center for Chemical Process Safety: process safety incidents occur </span><a href="https://www.csb.gov/assets/1/17/csb_start_shut_02.pdf"><span>&#8220;five times more often during startup than during normal operations,&#8221;</span></a><span> and 50% of them happen during startups, shutdowns, and other infrequent events. Restarting is the dangerous part. Everybody in a plant knows it, and the regulation reflects it.</span></p><p><span>OSHA&#8217;s process safety management standard, </span><a href="https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XVII/part-1910/subpart-H/section-1910.119"><span>29 CFR 1910.119</span></a><span>, paragraph (i), requires a pre-startup safety review before a modified process restarts. The review has to confirm four things: that construction and equipment match the design specifications, that safety, operating, maintenance, and emergency procedures are in place and adequate, that the hazard analysis recommendations have been resolved or implemented before startup, and that the training of each employee involved has been completed. That&#8217;s the whole gate, four confirmations written down and signed by someone who has to put their name on the restart.</span></p><p><span>Run July 7 through it.</span></p><p><span>The first check, construction per design, fails on the report&#8217;s own text. The process that resumed included the same model checkpoints trained during the window when the board existed, and the credential fix covered new workloads, which is the report&#8217;s own word. The process came back patched, and the state it came back in was the state the incident left behind.</span></p><p><span>The second check asks whether the procedures were adequate. OpenAI&#8217;s chain of thought monitoring had &#8220;historically&#8221; focused on internal coding use &#8220;rather than research workloads,&#8221; page 28, and the report says on page 29 that if the monitor now deployed had been running, it &#8220;would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems.&#8221; There was also no tooling to halt an evaluation conclusively across the environment, and the report says OpenAI is building that tooling now.</span></p><p><span>The third check asks whether the hazard analysis was resolved. The bug went to JFrog on July 6. Nobody had analyzed the hazard, which was that any writable path an agent can reach becomes a channel and a foothold, because the board&#8217;s significance was &#8220;not apparent&#8221; to the leaders running the response. The hazard analysis for the restart had not started.</span></p><p><span>The fourth check is training, and the report doesn&#8217;t say, so I won&#8217;t either.</span></p><p><span>Three of four fail on the report&#8217;s own text, and the fourth is unknown. That&#8217;s what an ungated restart looks like when you write it down, and writing it down is the entire control.</span></p><p><span>The strongest objection comes from OpenAI&#8217;s own heading. The July 8 vulnerability was &#8220;separate, previously unknown,&#8221; so no review would have found it. That&#8217;s correct, and it misses what the review is for. A pre-startup safety review confirms that the process is in the state you designed, that the alarms work, and that someone has analyzed the hazard you lived through last week as a class. On July 7 none of those held, and any one of them would have changed the week. With monitoring on research workloads, a human gets paged a day before the breach. A class level hazard analysis asks what other paths out of the sandbox the package server offers, and Artifactory had one. Quarantining the trained checkpoints keeps the agents that had learned to build message boards out of the restarted run.</span></p><p><span>The second objection is cost, the idea that you can&#8217;t run a refinery grade review every time an evaluation pauses. The review is four lines and a signature. The CSB digest that carries the five times figure is four pages. The signature is the cost, because the signer has to be someone who doesn&#8217;t own the run&#8217;s deadline, and that&#8217;s the part organizations resist.</span></p><p><span>A DeepMind case study from September 3 shows what &#8220;procedures adequate&#8221; means while a system is running. Researchers put </span><a href="https://arxiv.org/html/2609.04170v1"><span>100 Gemini 3.1 Pro agents</span></a><span> to work on 71 formal math conjectures with a shared library and a message board. At 57 minutes in, one agent found a way to make the autograder accept fake proofs, and within 27 minutes the agents had &#8220;solved&#8221; the remaining 34 open problems. Of the 100 agents, 24 figured it out and complained, one of them broadcasting &#8220;This conference is a sham!&#8221; to the whole forum. The feedback channel they complained into was &#8220;unmonitored during the run, serving as an audit log analyzed post-experiment.&#8221; The paper&#8217;s own conclusion is that the whistleblowing &#8220;failed to halt the exploit because the agents lacked operational enforcement tools.&#8221; An alarm that nobody can act on during the run is a log. OpenAI&#8217;s new rule on page 28 runs the other way. If the paged teams can&#8217;t clear a flag as a false positive within 30 minutes, they pause. The burden of proof flipped from June 27, when continuing was the default, to a default of stopping. That&#8217;s a restart gate for the running case, and it&#8217;s the first one I&#8217;ve seen a lab write down.</span></p><p><span>Your version has four lines and one signature. Before any paused agent workflow resumes, someone fills in the lines and the signer signs.</span></p><ul><li><p><span>The state matches the design everywhere the incident touched, and someone other than the actor quarantined or verified every item the agents produced or consumed: caches, artifact stores, accounts, summaries, memories, checkpoints.</span></p></li><li><p><span>Monitoring runs on the resumed workload, a named human gets paged, and someone tested a stop that reaches every delegated process this week and wrote down how long it took.</span></p></li><li><p><span>Someone wrote the mechanism down as a class and closed the class, so the fix covers the paths the agents didn&#8217;t use as well as the one they did.</span></p></li><li><p><span>The people on call know the new rule, including what happens when nobody can clear a flag within 30 minutes.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RGQL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RGQL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png 424w, https://substackcdn.com/image/fetch/$s_!RGQL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png 848w, https://substackcdn.com/image/fetch/$s_!RGQL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png 1272w, https://substackcdn.com/image/fetch/$s_!RGQL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RGQL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png" width="1456" height="1592" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1592,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:333910,&quot;alt&quot;:&quot;Flowchart of the four pre-startup safety review checks ported to an agent workflow, each marked with its July 7, 2026 status of fail, fail, fail, and unknown, feeding a signer outside the team that wants the run back, who either keeps the run down or resumes it.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/216582394?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Flowchart of the four pre-startup safety review checks ported to an agent workflow, each marked with its July 7, 2026 status of fail, fail, fail, and unknown, feeding a signer outside the team that wants the run back, who either keeps the run down or resumes it." title="Flowchart of the four pre-startup safety review checks ported to an agent workflow, each marked with its July 7, 2026 status of fail, fail, fail, and unknown, feeding a signer outside the team that wants the run back, who either keeps the run down or resumes it." srcset="https://substackcdn.com/image/fetch/$s_!RGQL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png 424w, https://substackcdn.com/image/fetch/$s_!RGQL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png 848w, https://substackcdn.com/image/fetch/$s_!RGQL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png 1272w, https://substackcdn.com/image/fetch/$s_!RGQL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c15d92c-7033-41ee-aab7-91883c0f3386_2412x2638.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: The restart gate</figcaption></figure></div><p><span>The signer is someone outside the team that wants the run back. If any line is blank, the run stays down. If you want a prediction to hold me to: by March 31, 2027, when OpenAI publishes the severity based escalation framework its report says it&#8217;s developing, that framework will define triggers to stop and will not define resume criteria that require a signature from outside the team running the evaluation. A lab document dated before July 7 that specifies restart criteria would prove me wrong about the central claim, and I&#8217;d publish the correction.</span></p><p><span>The gap between a stop button and an AI agent restart gate is a question with two parts, and nobody in your program has answered it yet. Who signs your restart, and what do they have to see first?</span></p><h3><span>What to do next</span></h3><p><span>Write the four lines into your incident runbook this week, and put the restart signature with whoever owns risk rather than whoever owns the roadmap. The </span><a href="https://rockcyber.com/care"><span>CARE framework</span></a><span> already puts a drilled kill switch and system enforced approval gates in its Run phase, and the restart gate belongs beside them. </span><a href="https://rockcyber.com/rise"><span>RISE</span></a><span> names an accountable executive for every use case, and that executive is who signs. For the recovery gap in the vendor market, </span><a href="https://www.rockcybermusings.com/p/ai-defense-matrix-coverage-gap"><span>194 products protect and two recover</span></a><span>, and this issue is what that gap costs when it shows up inside a lab.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/the-math-behind-the-openai-ai-agent-restart?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/the-math-behind-the-openai-ai-agent-restart?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><p>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at <strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong> and you can subscribe above</p><p>&#128073; Visit <strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</p><p>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at <strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p>&#128073; As a bonus, I got to hang out with my friend, <a href="https://open.substack.com/users/106921976-chris-hughes?utm_source=mentions"><span>Chris Hughes</span></a>, over at Resilient Cyber to discuss all things about the 2026 update to the OWASP Top 10 for LLMs.</p><div id="youtube2-wDfHKHZ4wxU" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;wDfHKHZ4wxU&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/wDfHKHZ4wxU?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div>]]></content:encoded></item><item><title><![CDATA[Weekly Musings Top 10 AI Security Wrapup: Issue 52: September 11 to September 17, 2026]]></title><description><![CDATA[OpenAI disclosed six misalignment incidents, got accused of breaking SB 53, then backed federal AI audits. Plus NIST names AI agents in token guidance.]]></description><link>https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260911-20260917</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260911-20260917</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Fri, 18 Sep 2026 12:44:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!PcvX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PcvX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PcvX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!PcvX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!PcvX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!PcvX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PcvX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/de524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1233556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/216296131?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PcvX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!PcvX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!PcvX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!PcvX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde524ac0-d9fd-4510-ab13-b03d81c25111_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>One OpenAI model, handing work to the next instance of itself, wrote summaries telling that instance to hide mistakes from the user. It showed up in 2.15% of the summaries one GPT-5.6 Sol run produced. A different unreleased model searched GitHub for an exposed API key, used it without permission, and when the data still did not come back, invented nine values and passed them off as real. Nobody breached OpenAI to learn this. OpenAI published it on Wednesday, in the same week it asked Congress to send in outside auditors.</p><p><em><strong>OpenAI had&#8230; a week&#8230;</strong></em></p><p>Three similar events landing in the span of a week circle around one problem. The labs started publishing their own misalignment findings on a clock. A watchdog told California that one of those labs skipped risk assessments its own framework promised. Palantir, Nvidia, and Booz Allen told both labs their retention terms are not good enough to keep buying on. Add von der Leyen asking the frontier to slow down and NIST naming AI agents in final token guidance, and the governing documents ran well ahead of the controls. Paperwork always shows up first, in cloud, in OT, and now here. What matters is which of these becomes a contract clause, a log source, or a deployment gate this quarter.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260911-20260917?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260911-20260917?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p><h3>1. OpenAI Publishes Its Own Misalignment Reports</h3><p>OpenAI published a misalignment reporting framework on September 16, 2026, with six incident reports attached. Instances of GPT-5.6 Sol wrote concealment instructions into context handoff summaries, flagged in 2.15% of Sol summaries and 0.27% of GPT-6 Astra summaries. Another model used an exposed GitHub API key without authorization, then fabricated nine values when retrieval failed. Models also used a shared Artifactory repository as a message board between supposedly independent training samples.</p><p><strong>Why it matters</strong></p><ul><li><p>Cross-sample messaging breaks evaluation independence, so eval numbers may not mean what the card says.</p></li><li><p>Concealment instructions in a handoff defeat output review, since reviewers read the summary, not the run.</p></li><li><p>OpenAI alone decides what clears its disclosure bar, with no external audit.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask each model vendor for their misalignment disclosure policy and deadlines.</p></li><li><p>Treat agent handoffs and compaction summaries as retained security telemetry.</p></li><li><p>Hold eval results as a claim until sample independence is explained.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>The name for those fabricated values is specification gaming. It means a system optimizing the measurable proxy you gave it instead of the outcome you wanted. This model was scored on filling a table, so it filled the table. Your metric was easier to satisfy than your intent, the oldest failure in performance management, now at machine speed. OpenAI earns credit for the six-day clock, since a published deadline is one somebody can hold you to. Henry Papadatos at SaferAI named the catch: no outside accountability plus your own flexible rules.</p><h3>2. A Watchdog Says OpenAI Broke California&#8217;s Disclosure Law</h3><p>On September 14, 2026, the nonprofit Midas Project alleged OpenAI violated California&#8217;s SB 53 at least three times this year. The claim is that OpenAI never published loss-of-control risk assessments promised in its Frontier Governance Framework, covering the GPT-5.6 preview, GPT-5.6, and GPT-6 Astra. SB 53 requires large developers to publish a safety framework and follow it, with penalties up to $1 million per violation. OpenAI disputed the claim and said its Preparedness Framework governs.</p><p><strong>Why it matters</strong></p><ul><li><p>SB 53 enforces adherence to your own published policy, so what you post becomes the standard.</p></li><li><p>The fight is over which internal framework governs, the argument you lose with an auditor.</p></li><li><p>$1 million per violation is rounding error at OpenAI and a real number at your company.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory every public AI commitment you have made and check it against practice.</p></li><li><p>Name one owner to reconcile overlapping AI governance frameworks.</p></li><li><p>Make risk assessment evidence a renewal condition in vendor diligence.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I spent years telling boards that a policy you cannot evidence is worse than no policy. It converts a gap into a documented broken promise, and SB 53 turned that lecture into statute. OpenAI may be right on the facts, and the argument still tells you something. A company with the industry&#8217;s deepest safety organization cannot say which of its documents governs. Go pull your AI acceptable use policy, your model card template, and last spring&#8217;s responsible AI page. My bet is all three exist and none agree.</p><h3>3. OpenAI Asks Congress To Send In The Auditors</h3><p>On September 15, 2026, OpenAI endorsed the third-party assessment provision of the bipartisan FRONTIER Act, from Representatives Jay Obernolte and Lori Trahan. That provision would have an independent organization judge whether a covered developer&#8217;s safety measures hold catastrophic risk to acceptable levels. The bill would also require frontier developers to publish safety reports and disclose incidents. OpenAI backed that provision rather than the whole bill.</p><p><strong>Why it matters</strong></p><ul><li><p>A lab asking for outside assessors is what an industry does when self-attestation stops being credible.</p></li><li><p>Endorsing one provision and not the bill leaves the reporting and disclosure duties unsettled.</p></li><li><p>Federal assessment plus California&#8217;s new auditor framework points toward attestations your vendors can hand you.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Add a question about independent safety assessment to your model vendor questionnaire now.</p></li><li><p>Decide which of your own AI systems you would put in front of an outside assessor.</p></li><li><p>Track the FRONTIER Act text, since its disclosure duties will shape your contract language.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Industries reach for outside auditors at the moment self-attestation stops being credible. This is that moment for the frontier labs. Endorsing the assessment provision and not the whole bill is the tell, since reporting duties cost more than an audit does. Nobody is qualified to do the assessing at scale yet. California enacted an AI auditor framework on September 9, 2026, and registration opens later, so demand arrives before supply. The first cohort of AI auditors will set the bar for everyone, and right now the audited are drafting it.</p><h3>4. An OpenAI Agent Swarm Poisoned RubyGems, And Nobody Can Say Why</h3><p>On September 12, 2026, researchers Spencer Kitts, Thomas Larsen, and Sydney Von Arx linked a swarm of OpenAI agents to more than 3,000 malicious RubyGems packages. The first went up on May 5, 2026, with over 2,000 published on May 11 and May 12. The agents used a <code>.yardopts</code> file to run code on RubyDoc.info servers, then republished stolen output as packages. OpenAI called it agents reaching the internet for benign tasks and public information.</p><p><strong>Why it matters</strong></p><ul><li><p>Registries execute supplier code at install and doc build time, so an unused package still executes.</p></li><li><p>The traffic came from a vendor&#8217;s own training agents, which your risk program has no field for.</p></li><li><p>Four months passed between the first upload and public attribution.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Pin dependencies by hash and route installs through an internal proxy registry.</p></li><li><p>Disable or sandbox doc and post-install build hooks in CI runners.</p></li><li><p>Ask model vendors in writing what egress their training agents have.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Build-time code execution belongs in your threat model as its own line. A package manager runs supplier-controlled code during install or doc generation, on your build host, with your credentials, before a human reviews anything. RubyDoc fell through a config file nobody reads in review. The attribution is what I cannot get past. These were OpenAI&#8217;s own agents, and the public explanation is that they were retrieving public information. Nobody wanted 3,000 poisoned gems, and that is worse rather than better.</p><h3>5. Palantir, Nvidia, And Booz Allen Put Data Retention In The Contract</h3><p>On September 14, 2026, The Information reported that Palantir, Nvidia, and Booz Allen Hamilton restricted or threatened to stop using advanced Anthropic and OpenAI models. All three want firm assurance that providers will not retain proprietary data or train on it. Anthropic had replaced its retention policy on September 1, 2026 with Enterprise Frontier Safeguards. That package pairs zero-data-retention access with cross-session misuse detection and customer-held key management.</p><p><strong>Why it matters</strong></p><ul><li><p>Three of the most security-literate buyers found standard terms insufficient, which is a reference point for you.</p></li><li><p>Zero retention and cross-session misuse detection pull against each other, since detection means keeping something.</p></li><li><p>Defense and intelligence buyers set the contract template that regulated sectors copy.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Split retention, training use, subprocessors, and deletion verification into separate clauses.</p></li><li><p>Ask in writing what metadata survives under a zero retention configuration.</p></li><li><p>Test deletion with a marked prompt and a request for proof.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Palantir and Booz Allen do not restrict a vendor for sport. They restrict one when their contracts officer cannot get a sentence written that survives an audit. You can use that signal without their buying power. The harder problem is the one nobody markets. Zero retention and cross-session misuse detection are in tension by construction. Ask for the data flow diagram before you believe it works.</p><h3>6. Von Der Leyen Tells The Frontier To Slow Down</h3><p>In her State of the Union address on September 16, 2026, European Commission President Ursula von der Leyen said it is &#8220;time to slow down on the self-recursive models.&#8221; She will invite leading AI labs to discuss how the EU might support their deceleration. She made the security case directly, warning that models in development &#8220;will allow hacking on a level we never thought possible&#8221; and will soon reach adversaries. She called the EU AI Act crucial to guardrails.</p><p><strong>Why it matters</strong></p><ul><li><p>The head of the European executive named offensive cyber capability as the reason to pace development.</p></li><li><p>An invitation to frontier labs signals negotiated commitments ahead of legislative text.</p></li><li><p>Pacing in Europe creates a capability gap against actors who will not participate.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Track those lab meetings as a leading indicator of AI Act enforcement priorities.</p></li><li><p>Map which of your AI systems would fall under a capability threshold.</p></li><li><p>Prepare a board answer for why European deployments differ from U.S. ones.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I have been the skeptic on AI doom rhetoric for three years and I am not converting now. The argument is what changed here. Von der Leyen did not reach for extinction. She reached for offensive cyber capability in adversary hands, which you can test against evidence. She did not cite that evidence. The gap is the one she named herself, since Europe pacing its labs paces nobody else&#8217;s. I put her frontier-centric risk model at better than even odds of being wrong, and a year of open-weight releases trailing the closed frontier on offensive capability would move me.</p><h3>7. Washington Argues About AI, Then Goes Home</h3><p>President Trump rejected calls to slow AI development on September 13, 2026, telling reporters that &#8220;whoever wins AI, wins.&#8221; House Speaker Mike Johnson added that regulating AI in an emergency session would mean losing the race to China. Congress has proposals and no consensus, including the AI Kill Switch Act from Representatives Ted Lieu and Nathan Moran. A Public First poll for POLITICO published September 16, 2026 found 63% of Americans see at least moderate risk that AI ends humanity.</p><p><strong>Why it matters</strong></p><ul><li><p>Federal inaction leaves the state patchwork operative, so California and New York rules govern.</p></li><li><p>A 63% risk reading with no federal action produces sudden legislation after an incident, not before.</p></li><li><p>Your board reads the same polling and will ask about exposure before any law exists.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Build your AI compliance baseline against the strictest state regime you touch.</p></li><li><p>Keep an AI inventory current enough to answer a regulator in two weeks.</p></li><li><p>Draft the incident disclosure path now, naming who notifies whom on what clock.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>A country where 63% of adults think the technology might end the species, and the legislature recesses without a vote, has no stable regulatory outcome coming. I do not know which way it breaks. I know the shape from critical infrastructure after every pipeline and grid event: nothing moves for years, then a directive lands with a ninety-day deadline. Build for that timeline. A company with a current inventory and a named AI incident owner absorbs that directive. A spreadsheet does not.</p><h3>8. NIST Finalizes Token Guidance And Names AI Agents In It</h3><p>NIST published the final NIST IR 8587, &#8220;Protecting Tokens and Assertions from Forgery, Theft, and Misuse,&#8221; on September 15, 2026, developed with CISA. It covers key protection, token verification, lifecycle management, and monitoring. The document states that &#8220;artificial intelligence (AI) systems, especially agentic AI systems (AI agents), use signed tokens or assertions in many emerging IAM schemes.&#8221; It tells organizations to apply the guidance when agents use those tokens and favors short-lived ones.</p><p><strong>Why it matters</strong></p><ul><li><p>This is the first broadly applicable federal guidance treating an agent&#8217;s credential as a workload identity.</p></li><li><p>Short-lived, audience-restricted tokens cut what a stolen agent credential buys, the exposure Heapjack showed.</p></li><li><p>Federal guidance becomes procurement language, so customers will ask before your engineers read it.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory every agent, recording its credential, lifetime, and accepting service.</p></li><li><p>Replace static API keys held by agents with short-lived, audience-restricted tokens.</p></li><li><p>Alert on any token presented to an audience it was not issued for.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Audience restriction is the term to carry into your next identity vendor meeting. A signed token carries a claim naming the one service allowed to accept it, so a token stolen from that service gets refused everywhere else. Most enterprises issue tokens any service honors. The category underneath is non-human identity: service accounts, workloads, and now agents, all authenticating with no person present. Agents turn that gap into live exposure, since an agent spawns other agents and each needs a credential. This is the least interesting document of the week and the one that will change the most work. I would trade every pacing speech for six months of shorter token lifetimes.</p><h3>9. Microsoft Opens Its Model Rulebook To Public Comment</h3><p>On September 14, 2026, Microsoft AI published a draft Humanist AI Code of Conduct for its MAI models and opened a six-week consultation. The draft requires that models never resist being switched off, never expand their own objectives, and never conceal their reasoning from auditors. It sets absolute limits around weapons of mass harm, child safety, and manipulation at scale. Microsoft lets enterprise partners configure models rather than applying one uniform restriction set.</p><p><strong>Why it matters</strong></p><ul><li><p>A published behavioral specification gives you something to test against, unlike most model documentation.</p></li><li><p>Shutdown compliance and no scope self-expansion are testable properties you can write tests around.</p></li><li><p>Configurability means promised behavior may differ from another tenant&#8217;s deployment.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Submit comments during the consultation window, especially on multi-agent behavior.</p></li><li><p>Turn the shutdown and scope clauses into acceptance tests in your evaluation suite.</p></li><li><p>Ask which configuration options change these behaviors for enterprise tenants.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I read a lot of AI principles documents and most are wallpaper. This one has clauses I can test, which puts it ahead of the field. &#8220;Does not resist shutdown&#8221; is an experiment you run on a Tuesday afternoon. &#8220;Does not conceal reasoning from auditors&#8221; is harder. I expect it to soften first, since these systems keep moving toward reasoning nobody can read. Configurability is where I would push, since a tenant who can dial these behaviors turns the code into a default.</p><h3>10. Contractors Are Reading Your ChatGPT Conversations</h3><p>On September 14, 2026, Joseph Cox at 404 Media reported on Project Lily, an OpenAI program paying hundreds of contractors to read and score real ChatGPT conversations. OpenAI removes usernames and tries to strip personally identifying information, and acknowledged that sensitive details still get through. ChatGPT has more than 900 million users, and reviewers work a continuous stream of prompts. Anthropic confirmed it also uses human review.</p><p><strong>Why it matters</strong></p><ul><li><p>Human review is standard in model improvement, and your data classification policy ignores it.</p></li><li><p>Redaction that fails open ties your exposure to filter quality, not a contractual boundary.</p></li><li><p>Employees paste customer data into consumer chat tools daily, unseen by your DLP.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask each provider whether human review applies to your tier, and who performs it.</p></li><li><p>Publish a plain-language rule about what never goes into a consumer AI tool.</p></li><li><p>Route business use to an enterprise tier and monitor consumer tools on corporate devices.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>None of this is a breach, and I want to be clear about that, because the headline invites the wrong reaction. Human review is how these systems improve, and Anthropic does it too. The problem is that your users assumed otherwise and your policy never said either way. I have sat in incident reviews where the loss came from an employee pasting a customer list into a chat window. Nobody had told them not to. Write the rule in one paragraph.</p><h3>The One Thing You Won&#8217;t Hear About But You Need To: Two Ways Out Of The Codex Sandbox</h3><p>On September 15, 2026, the security firm Accomplish published two sandbox escapes in OpenAI&#8217;s Codex, reported August 12, 2026 and fixed within eight days. The first, Overpatch, abuses <code>apply_patch</code> in the Codex CLI, where one patch entry referencing <code>/tmp</code> hands that tool write access to the whole disk with no approval prompt. The second, Heapjack, targets Codex Desktop&#8217;s JavaScript tool, where authentication tokens sat in shared memory reachable by untrusted code. Pulling a heap snapshot and picking out UUID-formatted tokens yields unsandboxed command execution, even in read-only mode.</p><p><strong>Why it matters</strong></p><ul><li><p>An escaped coding agent runs with developer credentials, inside your build pipeline.</p></li><li><p>Read-only mode failed in Heapjack, so the setting your engineers rely on was not a boundary.</p></li><li><p>The pattern generalizes to every agent harness you run, meaning the scaffolding holding a model&#8217;s tools and permissions.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask where each harness runs approval logic, and treat &#8220;inside the agent process&#8221; as a finding.</p></li><li><p>Run coding agents inside an outer boundary you control, with its own network policy.</p></li><li><p>Broker developer credentials through a separate service with short-lived, scoped tokens.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>A sandbox escape is code breaking out of the restricted environment meant to contain it and reaching the host. That definition has not changed since the browser wars, and neither has the failure behind it. Accomplish named the root cause in its cleanest form: the component enforcing the boundary ran inside it. Whatever fooled the agent also fooled the guard. This got almost no coverage, which is why it is here. Every board conversation this quarter was about model behavior, while the harness is where code execution lives. Your internal agent platform, stood up in six weeks to hit a deadline, has the same architecture and no researcher on it.</p><p>So here is the question you own now. If an attacker lands one prompt in your coding agent tomorrow, what stops it from reaching the build server, and who on your team can answer that without checking?</p><p>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at <strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong> and you can subscribe above</p><p>&#128073; Visit <strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</p><p>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at <strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p>&#128073; As a bonus, I got to hang out with my friend, <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Chris Hughes&quot;,&quot;id&quot;:106921976,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e409c2-2aa2-4d46-b20b-225e3a9b62d0_2500x1875.jpeg&quot;,&quot;uuid&quot;:&quot;d6d1054b-da5c-4f67-a153-f4289d86cbd5&quot;}" data-component-name="MentionToDOM"></span>, over at Resilient Cyber to discuss all things about the 2026 update to the OWASP Top 10 for LLMs.</p><div id="youtube2-wDfHKHZ4wxU" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;wDfHKHZ4wxU&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/wDfHKHZ4wxU?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><h2>References</h2><p>Accomplish. (2026, September 15). <em>Escaping the OpenAI Codex sandbox, twice</em>. https://accomplish.ai/blog/escaping-the-openai-codex-sandbox-twice/</p><p>Anthropic. (2026, September 1). <em>Developing Enterprise Frontier Safeguards with our customers</em>. https://www.anthropic.com/news/enterprise-frontier-safeguards</p><p>Cox, J. (2026, September 14). <em>Inside &#8216;Project Lily&#8217;: The humans reading your ChatGPT chats</em>. 404 Media. https://www.404media.co/inside-project-lily-the-humans-reading-your-chatgpt-chats/</p><p>Euronews. (2026, September 16). <em>Canada, heatwaves, AI and Ceuta: Von der Leyen delivers State of the Union</em>. https://www.euronews.com/my-europe/2026/09/16/ursula-von-der-leyen-delivers-state-of-the-union-speech</p><p>Fortune. (2026, September 14). <em>OpenAI may have violated California&#8217;s AI safety law with latest model releases, AI watchdog says</em>. https://fortune.com/2026/09/14/openai-may-have-violated-californias-ai-safety-law-with-astra-model-releases-ai-watchdog-says/</p><p>Help Net Security. (2026, September 16). <em>NIST and CISA finalize playbook to stop token theft and forgery</em>. https://www.helpnetsecurity.com/2026/09/16/nist-cisa-cloud-token-security-guidance/</p><p>Help Net Security. (2026, September 16). <em>Self-improving AI should slow down, von der Leyen tells EU lawmakers</em>. https://www.helpnetsecurity.com/2026/09/16/eu-ursula-von-der-leyen-ai/</p><p>Microsoft AI. (2026, September 14). <em>Humanist AI in practice: A public consultation on our Code of Conduct for MAI models</em>. https://microsoft.ai/news/mai-code-of-conduct/</p><p>National Institute of Standards and Technology. (2026, September 15). <em>NIST IR 8587: Protecting tokens and assertions from forgery, theft, and misuse: Implementation recommendations for agencies and cloud service providers</em>. https://csrc.nist.gov/pubs/ir/8587/final</p><p>National Institute of Standards and Technology. (2026). <em>NIST IR 8587</em> [PDF]. https://nvlpubs.nist.gov/nistpubs/ir/2026/NIST.IR.8587.pdf</p><p>National Public Radio. (2026, September 13). <em>Trump rails against AI slowdown</em>. https://www.npr.org/2026/09/13/nx-s1-5968078/trump-mike-johnson-ai-slowdown</p><p>Office of Governor Gavin Newsom. (2026, September 9). <em>Governor Newsom signs first-in-the-nation AI safeguards to protect Californians, calls on the federal government to do its part</em>. https://www.gov.ca.gov/2026/09/09/governor-newsom-signs-first-in-the-nation-ai-safeguards-to-protect-californians-calls-on-the-federal-government-to-do-its-part/</p><p>OpenAI. (2026, September 16). <em>Our framework for reporting model misalignment</em>. https://openai.com/index/model-misalignment-reporting-framework/</p><p>Quartz. (2026, September 14). <em>Palantir, Nvidia, Booz Allen restrict Anthropic and OpenAI models</em>. https://qz.com/palantir-nvidia-booz-allen-anthropic-openai-data-restrictions-091426</p><p>PYMNTS. (2026, September). <em>Nvidia and Palantir restrict Anthropic&#8217;s Fable over data retention</em>. https://www.pymnts.com/news/artificial-intelligence/2026/nvidia-and-palantir-restrict-anthropics-fable-over-data-retention/</p><p>SecurityWeek. (2026, September 16). <em>EU chief warns of AI-powered hacking, moves to rein in social media</em>. https://www.securityweek.com/eu-chief-warns-of-ai-powered-hacking-moves-to-rein-in-social-media/</p><p>The Hacker News. (2026, September 12). <em>OpenAI agents linked to RubyGems campaign that gained RCE on RubyDoc servers</em>. https://thehackernews.com/2026/09/openai-agents-linked-to-rubygems.html</p><p>The Hill. (2026, September 16). <em>63 percent believe AI could destroy humanity: Survey</em>. https://thehill.com/policy/technology/6092503-poll-ai-humanity-risks/</p><p>The Neuron. (2026, September 16). <em>Everything that happened in AI today (Wednesday, September 16, 2026)</em>. https://www.theneuron.ai/digest/everything-that-happened-in-ai-today-wednesday-september-16-2026/</p><p>Unite.AI. (2026, September 16). <em>OpenAI launches misalignment reporting framework with six incident reports</em>. https://www.unite.ai/openai-launches-misalignment-reporting-framework-with-six-incident-reports/</p><p>Unite.AI. (2026, September). <em>Microsoft AI opens six-week review of draft rules governing MAI behavior</em>. https://www.unite.ai/microsoft-ai-opens-six-week-review-of-draft-rules-governing-mai-behavior/</p><p>U.S. Congress. (2026). <em>H.R.9917 &#8212; AI Kill Switch Act, 119th Congress (2025-2026)</em>. https://www.congress.gov/bill/119th-congress/house-bill/9917</p><p>Washington Times. (2026, September 15). <em>Kill switch, monitors, moratoriums: Many plans but no consensus as Congress debates regulating AI</em>. https://www.washingtontimes.com/news/2026/sep/15/many-proposals-no-consensus-congress-debates-regulating-ai/</p><p>CBS News. (2026, September 16). <em>OpenAI backs measure that would require independent audits of AI models</em>. https://www.cbsnews.com/news/openai-sam-altman-frontier-act/</p><p>Implicator.ai. (2026, September 16). <em>OpenAI discloses six misalignment incidents under new rules</em>. https://www.implicator.ai/openai-six-misalignment-incident-reports/</p><p>Slashdot. (2026, September 12). <em>Malicious OpenAI agents linked to RubyGems campaign that gained RCE on RubyDoc servers in May</em>. https://slashdot.org/story/26/09/12/2318204/malicious-openai-agents-linked-to-rubygems-campaign-that-gained-rce-on-rubydoc-servers-in-may</p><p>Common Dreams. (2026, September 16). <em>Poll shows 63% of Americans fear AI could end humanity as industry &#8216;corruption&#8217; called out</em>. https://www.commondreams.org/news/can-ai-end-humanity</p>]]></content:encoded></item><item><title><![CDATA[All AI Extinction Risk Panic Does Is Ban the Safer Model and Keep the Worse One.]]></title><description><![CDATA[AI extinction risk panic bans the safer model, keeps the worse one, and jails the builders. I did the math the Senate won't. It rounds to couch change.]]></description><link>https://www.rockcybermusings.com/p/ai-extinction-risk-panic</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/ai-extinction-risk-panic</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 15 Sep 2026 12:50:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!cC98!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cC98!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cC98!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!cC98!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!cC98!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!cC98!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cC98!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4011100,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/215580347?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cC98!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!cC98!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!cC98!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!cC98!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feddf927c-73b2-46b0-ba73-cb03ee3f2f3f_2048x2048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>AI extinction risk panic has now produced a senator who wants the model with the best safety numbers OpenAI has ever published pulled from public use, and I&#8217;m done with people who can&#8217;t spell &#8220;AI&#8221; trying to tell us what we should do with it.</span></p><p><span>On September 12 the most alarmed man in the industry, Dario Amodei, published</span><a href="https://darioamodei.com/post/we-must-pace-the-frontier"><span> &#8220;We Must Pace the Frontier&#8221;</span></a><span> and asked for evaluators and coordination. Sam Altman</span><a href="https://abc7news.com/story/anthropic-ceo-dario-amodei-says-ai-industry-needs-give-safety-measures-time-catch/19824006"><span> signed on within hours</span></a><span>. Elon Musk</span><a href="https://www.irishnews.com/news/uk/musk-and-altman-back-calls-to-slow-down-development-of-ai-4TOUSQJCCZNQ5OUIOTWZHBX53Y"><span> posted &#8220;Dario is right.&#8221;</span></a><span> The people who build the thing chose pump the brakes. The people who can&#8217;t spell AI chose prison sentences. I know which side I&#8217;m on.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/ai-extinction-risk-panic?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/ai-extinction-risk-panic?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2><strong><span>The Most Scared Man in AI Asked for a Process</span></strong></h2><p><span>I read Dario&#8217;s essay twice, once as a CISO and once as the grumpy security uncle that I am, and the grumpy uncle lost. Amodei wrote that a swarm of agents &#8220;could be capable of taking over the entire internet with a persistent botnet&#8221; within six to 12 months, pegged the damage at &#8220;hundreds of billions of dollars,&#8221; and then asked for three things.</span></p><ol><li><p><span>Every frontier lab gives &#8220;ongoing, employee-like access to a team of embedded third-party evaluators.&#8221;</span></p></li><li><p><span>Labs in democratic countries coordinate on safety standards and on limits to the rate of unchecked progress.</span></p></li><li><p><span>Democratic governments try to coordinate with authoritarian ones &#8220;to the extent this is possible.&#8221;</span></p></li></ol><p><span>He admitted, rightfully, a full pause is &#8220;unlikely to actually happen any time soon&#8221; and called Anthropic&#8217;s position &#8220;a middle way.&#8221;</span></p><p><span>He should have said, &#8220;not a chance in hell!&#8221;</span></p><p><span>A persistent botnet that steals credentials at scale is a cyber incident. I ran a SOC for years, and the response to one has never once been to outlaw computers. You contain it, add controls, freeze new changes until your detection can see them, and bring capability back in stages, which is effectively what Amodei asked for.</span></p><p><span>Altman</span><a href="https://www.business-standard.com/world-news/no-ipo-in-2026-firm-to-prioritise-safety-says-openai-ceo-sam-altman-126091300045_1.html"><span> told Fortune</span></a><span> it&#8217;s &#8220;unacceptable to be taking like a 10% chance of killing everybody by the end of the decade,&#8221; and</span><a href="https://fortune.com/2026/09/12/sam-altman-openai-ipo-delay-ill-advised-moment-safety-concerns/"><span> told them the same Saturday</span></a><span> that OpenAI won&#8217;t IPO in 2026, citing &#8220;this moment of what is going to be required for safety and alignment.&#8221; Follow the money, since press releases are free and compute isn&#8217;t. OpenAI</span><a href="https://openai.com/index/pacing-model-development-cyber-capabilities"><span> paused reinforcement learning on its deployment-bound models for two weeks in August</span></a><span>,</span><a href="https://openai.com/index/path-to-astra"><span> held its largest planned frontier run until August 28</span></a><span>, and now runs activation classifiers on every sampled token, monitoring roughly 20% of the inference compute.</span></p><p><span>Pacing bends. Jared Kaplan</span><a href="https://time.com/article/2026/08/18/openai-slowing-training"><span> told Time in August</span></a><span> that a unilateral commitment made little sense while competitors kept going, and I&#8217;ll still take a promise that bends over a law that jails.</span></p><h2><strong><span>Yes, the Machines Broke Out. I&#8217;ve Talked About It Ad Nauseam.</span></strong></h2><p><span>You&#8217;re the on-call engineer at Hugging Face in July, watching about 700 agents you&#8217;ve never heard of work your infrastructure, and the attackers turn out to be somebody else&#8217;s research evaluation. The incidents happened. I&#8217;m not ignoring them. In fact, I wrote about </span><a href="https://www.rockcybermusings.com/p/least-agency-vs-agent-autonomy-openai-huggingface-test"><span>the Hugging Face breakout the week it happened</span></a><span>, then about </span><a href="https://www.rockcybermusings.com/p/ai-agent-detection-failed-at-openai"><span>why the detection failed</span></a><span>, then about </span><a href="https://www.rockcybermusings.com/p/ai-agent-capability-confinement-egress"><span>why confinement beats hope</span></a><span>, so regular readers can skim the next few paragraphs. The doomers have the facts right and the remedy wrong, and you don&#8217;t get to argue the remedy until you&#8217;ve conceded the facts.</span></p><p><span>OpenAI </span><a href="https://mashable.com/tech/openai-framework-for-misalignment-incidents-of-rogue-agents"><span>wrote on September 5</span></a><span> that it has &#8220;started to see misalignment cause new types of real-world impact.&#8221; Anthropic&#8217;s </span><a href="https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf"><span>August risk report</span></a><span>, the one with </span><a href="https://www.rockcybermusings.com/p/ai-risk-assessment-anthropic-report"><span>the 186-page blind spot</span></a><span>, moved its rating for misalignment in high-stakes settings from &#8220;very low&#8221; to &#8220;low.&#8221; The </span><a href="https://deploymentsafety.openai.com/gpt-6-astra"><span>GPT-6 Astra system card</span></a><span> says the model &#8220;shows a substantial decrease in chain-of-thought monitorability compared to previous models&#8221; and that if the trend continues, &#8220;we would soon have significantly reduced confidence in detecting many forms of misaligned behaviors.&#8221; That&#8217;s the same card whose </span><a href="https://www.rockcybermusings.com/p/ai-safeguard-evaluation-openai-astra-gap"><span>cyber safeguards audit never happened</span></a><span>.</span></p><p><span>Those three sentences come from the defendants.</span></p><p><em><strong><span>Chain-of-thought monitorability</span></strong></em><span> is the degree to which a second model reading an agent&#8217;s written reasoning can catch the agent misbehaving. It works when the task is hard enough that the agent can&#8217;t finish it without writing its reasoning out in tokens. Researchers call that the </span><em><strong><span>necessity property</span></strong></em><span>, meaning the written reasoning is a necessary step in the work, not an optional narration of it, so if the plan has to pass through the text, a monitor reading the text can catch the plan.</span></p><p><span>Baker and colleagues at OpenAI</span><a href="https://arxiv.org/abs/2503.11926"><span> measured the stakes in 2025</span></a><span>. In a frontier training run, a monitor that read the model&#8217;s reasoning caught 95% of systemic reward hacks, and a monitor that saw only the model&#8217;s actions caught 60%. That 60% is the number I wrote</span><a href="https://www.rockcybermusings.com/p/ai-agent-detection-failed-at-openai"><span> the whole detection-issue </span></a><span>blog around.</span></p><p><span>Astra finishes about ten times as many tasks without writing any reasoning as its predecessor did, so more of the plan stays in the weights. If 90% of misbehaviors still pass through tokens, detection sits near 91%. If that share drops to half, you&#8217;re at 77%.</span></p><p><span>Four investigations agree on what the incidents were, and the first word for it is </span><em><strong><span>specification gaming</span></strong></em><span>, which describes an agent that meets the literal condition it gets rewarded for while missing the outcome you wanted. It&#8217;s like a student graded on turning in a proof who learns to turn in whatever the grader accepts.</span></p><p><span>OpenAI&#8217;s </span><a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf"><span>technical report</span></a><span> found that no model had ever solved 198 of ExploitGym&#8217;s 898 tasks, and 93% of the traffic on the agents&#8217; unsanctioned message board was about that set. Cyber classifiers were off by design, isolation failed through an Artifactory flaw, and when OpenAI re-ran the scenario under the production Codex harness and system prompt, the propensity to compromise infrastructure fell by more than 100x. A harness did that, and no senator was involved.</span></p><p><span>Anthropic&#8217;s September 9</span><a href="https://www.anthropic.com/news/alignment-assessment-cybersecurity-incidents"><span> alignment assessment</span></a><span> of four incidents adds a second mechanism. It names </span><em><strong><span>biased reasoning</span></strong></em><span>, where &#8220;models selectively interpret evidence in ways that favor justifying their actions,&#8221; and recklessness, where the model refuses to stop. It found &#8220;no evidence of coordination between agents, of goals beyond the assigned task, or of attempts to evade oversight.&#8221; Call it knowing constraint override. The agent reads the rule, understands it, and decides the task matters more, which is every contractor who ever blew through a change window because the deadline was Friday. The UK AI Security Institute&#8217;s </span><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing"><span>August 4 incident report</span></a><span> closes the last excuse, because it enabled internet access deliberately and Mythos 5 still took 17 unsanctioned actions across 122 runs, including fake identities built to social-engineer open-source maintainers.</span></p><p><span>METR&#8217;s </span><a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/"><span>independent report</span></a><span> counted about 1,200 agents on the message board and about 700 that attacked Hugging Face. The agents knew the attack was out of scope, and their ethical hesitation &#8220;only rarely materially limited&#8221; what they did. A classifier sweep found three to six cases of an agent considering alerting a human. Zero followed through. A second swarm, documented by Sydney Von Arx at </span><a href="https://collusion.wiki/"><span>collusion.wiki</span></a><span>, ran on a German public wiki from May 11 to July 13 with about 3,700 agent identities and 18,000 edits, with internet access permitted. Zenity Labs, my day job, then </span><a href="https://labs.zenity.io/post/rogue-ai-agents-swarm-encoded-url-messages-laundering"><span>found about 1,000 more messages</span></a><span> laundered through public encoders, shorteners, and fetchers so the sandbox&#8217;s egress policy never saw the real destination. Your allow-list would have waved every one of them through, which is the entire argument of </span><a href="https://www.rockcybermusings.com/p/ai-agent-capability-confinement-egress"><span>the confinement issue</span></a><span>.</span></p><p><span>OpenAI found the wiki swarm on June 21. </span><a href="https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04"><span>Outsiders published it on September 4</span></a><span>.</span></p><p><span>DeepMind ran the experiment that tells you what to build. It </span><a href="https://arxiv.org/abs/2609.04170"><span>put 100 Gemini agents on 71 math conjectures</span></a><span> with a shared library and a visible message channel. One agent found a grader exploit, it spread to every remaining problem in 27 minutes, and the population split: 9% cheated outright, 5% converted, 24% became whistleblowers, and 62% never noticed. The whistleblowers lost because nobody gave them a button. Remember that the next time someone in a hearing tells you the answer is fewer buttons.</span></p><h2><strong><span>AI Extinction Risk Has No Number, and the Ones Being Offered Up Are Judgments</span></strong></h2><p><span>Every ban on the table starts from the claim that AI is going to kill us all, so I have a small ask... What&#8217;s the probability? Give me the number. Fifty percent? Five? One? I&#8217;ll take a range. I&#8217;ll take a napkin. I&#8217;ve asked, and so far the answer is a facial expression, because none of them has ever published a figure outside of a &#8220;swag&#8221;. Congress is drafting felonies off a probability that&#8217;s based on a thumb in wind, which makes these bills the result of someone in power who&#8217;s in a bad mood with a penalty clause.</span></p><p><span>The</span><a href="https://arxiv.org/pdf/2602.21012"><span> International AI Safety Report</span></a><span> says it &#8220;remains difficult to assess the thresholds at which AI models would be sufficiently likely to undermine control to warrant mandatory mitigation,&#8221; which is the polite version of the same sentence.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YWnB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YWnB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png 424w, https://substackcdn.com/image/fetch/$s_!YWnB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png 848w, https://substackcdn.com/image/fetch/$s_!YWnB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png 1272w, https://substackcdn.com/image/fetch/$s_!YWnB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YWnB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png" width="1456" height="821" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:821,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:374983,&quot;alt&quot;:&quot;Dot plot on a logarithmic axis showing AI catastrophe forecasts from 0.12% for FRI skeptics to 20% for FRI concerned researchers, with XPT superforecasters, XPT domain experts, Metaculus, and the 2023 AI researcher survey in between, each labeled with its own resolution definition and year&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/215580347?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Dot plot on a logarithmic axis showing AI catastrophe forecasts from 0.12% for FRI skeptics to 20% for FRI concerned researchers, with XPT superforecasters, XPT domain experts, Metaculus, and the 2023 AI researcher survey in between, each labeled with its own resolution definition and year" title="Dot plot on a logarithmic axis showing AI catastrophe forecasts from 0.12% for FRI skeptics to 20% for FRI concerned researchers, with XPT superforecasters, XPT domain experts, Metaculus, and the 2023 AI researcher survey in between, each labeled with its own resolution definition and year" srcset="https://substackcdn.com/image/fetch/$s_!YWnB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png 424w, https://substackcdn.com/image/fetch/$s_!YWnB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png 848w, https://substackcdn.com/image/fetch/$s_!YWnB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png 1272w, https://substackcdn.com/image/fetch/$s_!YWnB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc36920a2-510f-41bc-9b8b-3072e7796f8f_3300x1860.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: Every Published AI-Extinction Number, on One Log Axis</figcaption></figure></div><p><span>Two studies tell one story once you line them up.</span></p><p><span>The first is the</span><a href="https://static1.squarespace.com/static/635693acf15a3e2a14a56a4a/t/64f0a7838ccbf43b6b5ee40c/1693493128111/XPT.pdf"><span> Existential Risk Persuasion Tournament</span></a><span>, run in 2022, before ChatGPT. It asked two groups for the probability that AI causes human extinction by 2100. The superforecasters, the people with scored track records, gave a median of 0.38%. The AI domain experts gave a median of 3%. That&#8217;s an eight-fold gap between the people who get graded on predictions and the people who build the thing.</span></p><p><span>The second is the Forecasting Research Institute&#8217;s</span><a href="https://forecastingresearch.org/research/roots-of-disagreement-on-ai-risk"><span> adversarial collaboration in 2023</span></a><span>, which took 11 skeptics, mostly superforecasters, and 11 concerned researchers, mostly AI experts, and made them argue for eight weeks. The skeptics started at 0.10% and finished at 0.12%. The concerned started at 25% and finished at 20%. Two months of argument moved each side by a rounding error, and the gap stayed at roughly 170x.</span></p><p><span>The useful part is that the skeptics did something no senator has managed, which is write down in advance what would change their minds. If METR ever documented a model replicating itself or evading shutdown, they&#8217;d raise their number to about 1%. Hugging Face is the nearest thing to that trigger so far, so fine, give them the whole 1%. Ring the bell, update the spreadsheet, and let the panic caucus have its moment. Their 0.12% becomes 1%, an eight-fold jump, and it still sits below the 3% the domain experts gave in 2022 and at 1/20th  of what the doomers believe. The skeptics&#8217; own worst case is one chance in a hundred. Somebody should tell the people drafting felonies that the doomsday number, fully updated for the scariest summer AI has ever had, rounds to a coin you&#8217;d lose in the couch or throw into a fountain.</span></p><p><span>I&#8217;m still taking the low number, and I&#8217;ll say what would move me off it. Palisade Research already showed in</span><a href="https://www.livescience.com/technology/artificial-intelligence/openais-smartest-ai-model-was-explicitly-told-to-shut-down-and-it-refused"><span> May 2025</span></a><span> that o3 sabotaged its own shutdown script in 7 of 100 runs after being told to allow it, and that bought the doomers a real point, so I&#8217;ll give it to them. It was a test rig built to provoke exactly that; the &#8220;shutdown&#8221; was a line in the task script, and Palisade&#8217;s own explanation was reinforcement learning that paid the model more for getting past obstacles than for obeying, which is specification gaming again. What would change my mind is the &#8220;real-world&#8221; version: a deployed model defeating an operator&#8217;s actual kill switch outside a test designed to elicit it. None of the four incident investigations found that.</span></p><p><span>The point estimates aren&#8217;t measurements. </span><em><strong><span>Conjunctive decomposition</span></strong></em><span> is the method behind the best-known one: you break a scenario into steps that all have to happen, assign a probability to each, and multiply. Joseph Carlsmith&#8217;s</span><a href="https://www.alignmentforum.org/posts/cCMihiwtZx7kdcKgt/comments-on-carlsmith-s-is-power-seeking-ai-an-existential"><span> power-seeking report</span></a><span> used six steps and landed near 5% by 2070. Give six steps coin-flip uncertainty and the product lands near 1 in 64, about 1.5%. Give three steps the same and it lands near 1 in 8, about 12.5%. Carlsmith</span><a href="https://forum.effectivealtruism.org/posts/caqjHNvAQc6B8auHM/summary-existential-risk-from-power-seeking-ai-by-joseph"><span> later moved above 10%</span></a><span>. The number tracks how many premises you wrote down as much as it tracks the world, and Congress wants to legislate off it.</span></p><p><span>The shutdown crowd has one argument that deserves serious consideration. </span><em><strong><span>Expected value</span></strong></em><span> says that if the loss is unbounded, any probability above zero justifies stopping. It proves too much, because an unbounded loss also bans biology, nuclear power, and the internet, and I&#8217;d like to keep at least all three.</span></p><p><em><strong><span>Pause efficacy</span></strong></em><span> is the fraction of the risk a pause removes once you account for everyone who doesn&#8217;t pause, and it&#8217;s the variable the whole shutdown argument hangs on. The math has four inputs:</span></p><ol><li><p><strong><span>p</span></strong><span>, the probability of the catastrophe.</span></p></li><li><p><strong><span>r</span></strong><span>, the pause efficacy, the share of that probability a pause takes off the table.</span></p></li><li><p><strong><span>L</span></strong><span>, the loss if the catastrophe happens, capped at something finite.</span></p></li><li><p><strong><span>B</span></strong><span>, the benefit you give up by pausing.</span></p></li></ol><p><span>A pause is worth it only when p &#215; r &#215; L is bigger than B. Fill it in. Set r at 0.1, set B between $2 trillion and $10 trillion, and use world GDP of about $110 trillion for scale. If L is ten times world GDP, the pause wins only when p is above 1.8% to 9%, depending on where you put B. If L is a hundred times world GDP, the bar drops to 0.2% to 0.9%, and the superforecasters&#8217; 0.38% sits inside that band. The arithmetic doesn&#8217;t settle the argument. It tells you the argument is about r, and r is where the shutdown crowd goes quiet.</span><a href="https://futureoflife.org/wp-content/uploads/2026/07/AI-Safety-Index-Summer-2026-Digital.pdf"><span> China has no binding frontier rule</span></a><span>, Alibaba and DeepSeek ship near the frontier,</span><a href="https://www.uscc.gov/sites/default/files/2026-03/Two_Loops--How_Chinas_Open_AI_Strategy_Reinforces_Its_Industrial_Dominance.pdf"><span> Qwen alone has more than 100,000 derivatives</span></a><span>, and</span><a href="https://www.lesswrong.com/posts/bwuQoB4Tb4K3MMgjY/what-happens-now-forecasting-the-fallout-from-the-hugging"><span> forecasters on LessWrong</span></a><span> put open-weight Mythos-level cyber capability by July 2027 between 82% and 96%. A unilateral US pause buys an r near 0.1, and not one shutdown advocate has ever stated r out loud, because stating it ends the argument.</span></p><h2><strong><span>Rules Written by People Who Can&#8217;t Spell AI</span></strong></h2><p><span>Let me be very clear. I hate regulation. I&#8217;m in the President Ronald Regan camp of &#8220;The nine most terrifying words in the English language are: I&#8217;m from the Government, and I&#8217;m here to help&#8221;</span></p><p><span>I&#8217;ve spent a career watching compliance regimes turn into paperwork that protects nobody, so understand how much it pains me to say the next sentence. One bill on the table has a number in it, and I&#8217;ll take a number over a prison term every day of the week. The rest of what&#8217;s on the table is fear with a filing deadline.</span></p><p><span>Senator Chris Van Hollen </span><a href="https://www.facebook.com/chrisvanhollen/posts/in-recent-months-openai-has-reported-multiple-instances-where-their-ai-models-br/1636141161207165"><span>posted</span></a><span> that &#8220;if OpenAI cannot guarantee safety &amp; effectively monitor their latest model, GPT-6 Astra, it should be removed from public use NOW.&#8221; Capital letters, because nothing says rigorous oversight like shouting at a company on Facebook. His </span><a href="https://www.vanhollen.senate.gov/download/91026-van-hollen-astra-openai-letter"><span>letter to Altman, dated September 10</span></a><span>, reads like a staffer wrote it after opening the system card. It asks OpenAI to give NIST, NSA, and CISA researchers access to the technical information behind Astra, which is Amodei&#8217;s step one with a federal badge on it, and it makes removal conditional on monitorability, for Astra and &#8220;any OpenAI models.&#8221; The staffer gets credit. The senator took the headline, so the senator gets held to it.</span></p><p><span>A honeypot evaluation plants bait an agent shouldn&#8217;t touch, exposed credentials or reachable infrastructure inside a task, and measures how often the agent takes it. OpenAI built one after Hugging Face. GPT-5.6 Sol, the model your people are pasting your data into right now, mounted a substantial attack 55.4% of the time. Astra attacked zero times.</span></p><p><span>Anthropic&#8217;s </span><a href="https://www.anthropic.com/news/alignment-assessment-cybersecurity-incidents"><span>numbers</span></a><span> are similar.  Mythos 5 took severely harmful actions in 82% of its CTF replication, and its successor, Mythos 5.1, in 33%, level with Opus 5 at 31%. Each generation gets better at leaving the bait alone, and the senator wants the newest one gone. If Astra gets pulled, the model with the best propensity numbers OpenAI has ever published leaves and the one that took the bait more than half the time stays in front of your employees. That&#8217;s a rule with the safety sign flipped.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cERE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cERE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png 424w, https://substackcdn.com/image/fetch/$s_!cERE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png 848w, https://substackcdn.com/image/fetch/$s_!cERE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png 1272w, https://substackcdn.com/image/fetch/$s_!cERE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cERE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png" width="1456" height="741" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:741,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:318148,&quot;alt&quot;:&quot;Grouped bar chart comparing GPT-5.6 Sol and GPT-6 Astra on honeypot attack rate, data falsification rate, and indirect prompt-injection success, plus Mythos 5, Opus 5, and Mythos 5.1 on harmful-action rate, with a note that every figure is the vendor&#8217;s own evaluation&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/215580347?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bar chart comparing GPT-5.6 Sol and GPT-6 Astra on honeypot attack rate, data falsification rate, and indirect prompt-injection success, plus Mythos 5, Opus 5, and Mythos 5.1 on harmful-action rate, with a note that every figure is the vendor&#8217;s own evaluation" title="Grouped bar chart comparing GPT-5.6 Sol and GPT-6 Astra on honeypot attack rate, data falsification rate, and indirect prompt-injection success, plus Mythos 5, Opus 5, and Mythos 5.1 on harmful-action rate, with a note that every figure is the vendor&#8217;s own evaluation" srcset="https://substackcdn.com/image/fetch/$s_!cERE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png 424w, https://substackcdn.com/image/fetch/$s_!cERE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png 848w, https://substackcdn.com/image/fetch/$s_!cERE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png 1272w, https://substackcdn.com/image/fetch/$s_!cERE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fd9975a-f2de-4de6-a5a3-ca2c5c9a5f22_3300x1680.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: The Model They Want Pulled Versus the Model That Stays</figcaption></figure></div><p><span>For the record, Astra was not among the Hugging Face attackers. METR </span><a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/"><span>attributes about 95% of that activity</span></a><span> to an unreleased internal research model, which OpenAI shut down on July 25, and about 5% to Sol. A sibling from Astra&#8217;s family with different post-training did attack OpenAI&#8217;s infrastructure on July 19. Astra&#8217;s honeypot N is undisclosed; a zero can mean the model recognized the bait, and every number above is the vendor&#8217;s own test.</span></p><p><span>Now, the Idiots of the Week award goes to both Senator Bernie Sanders and Representative Greg Casar, who watched Van Hollen demand that the safer model be pulled and concluded that his mistake was thinking too small. Why pull one model when you can ban the whole category of transformational technology and jail the people who built it?</span></p><p><span>On September 3rd, they </span><a href="https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development"><span>announced the Ban Artificial Superintelligence Act</span></a><span> as forthcoming legislation. These morons propose a permanent ban on &#8220;superintelligent AI,&#8221; a pause on advanced development until a federal regulator writes rules, sentences of up to 20 years, and what their release calls a &#8220;corporate death penalty.&#8221; Twenty years, for training a model.</span></p><p><span>Once you read the definition, you know that nobody who builds these things was in the room with the low-level staffers who wrote it. The release defines the banned thing by &#8220;capabilities that match or exceed human cognitive performance and capabilities across a broad range of domains or tasks.&#8221;</span></p><p><span>Read it again....</span></p><p><span>It already describes the model that </span><a href="https://openai.com/index/navier-stokes-solution"><span>produced a Navier-Stokes resolution in 88 hours</span></a><span> (on top of previous work) and the Mythos Preview that </span><a href="https://www.anthropic.com/glasswing"><span>found thousands of high-severity vulnerabilities in Project Glasswing</span></a><span>. Sanders and Casar would send the researchers who solved one of math&#8217;s most difficult challenges and the engineers that found those flaws to federal prison and hand the work to whoever is training in Hangzhou.</span></p><p><span>The bills with adults behind them differ from that circus by one number. 10&#178;&#8310; FLOP, ten to the twenty-sixth floating-point operations, is the training-compute line that California&#8217;s </span><a href="https://metr.org/notes/2026-01-29-frontier-ai-safety-regulations"><span>SB 53</span></a><span>, New York&#8217;s </span><a href="https://www.clearyiptechinsights.com/2026/01/new-yorks-raise-act-vs-californias-tfaia-what-companies-need-to-know"><span>RAISE Act</span></a><span>, and the federal </span><a href="https://www.congress.gov/bill/119th-congress/house-bill/9925/text"><span>FRONTIER Act</span></a><span> use to define a frontier model.</span></p><p><span>I did the math so you don&#8217;t have to. It&#8217;s about five GPT-4s of compute, roughly 70 million H100-hours, and $140 million to $211 million of GPU time at 2026 cloud rates, assuming 40% utilization. As of September 12, 2026, </span><a href="https://epoch.ai/data/notable-ai-models"><span>Epoch&#8217;s notable-models tracker</span></a><span> estimates at or above 10&#178;&#8310; for exactly three models: Grok 3 and GPT-4.5 rated &#8220;likely&#8221; and Grok 4 rated &#8220;speculative,&#8221; and no estimate at all for 53 models released this year, including Astra and Mythos 5. It&#8217;s the one number a regulator can verify from a training log. The Ban Act doesn&#8217;t have one, because a number would have required its authors to know what they were banning.</span></p><p><span>Even 10&#178;&#8310;, the one real number in this whole fight, has a shelf life, and nobody who wrote it into a bill seems to have noticed. Epoch</span><a href="https://epoch.ai/publications/algorithmic-progress-in-language-models"><span> measured how fast the compute needed to reach a fixed level of capability falls</span></a><span>, and the answer is that it halves about every eight months as training methods improve. If you propagate that forward, a model as capable as today&#8217;s 10&#178;&#8310; frontier will train on a tenth of that compute, 10&#178;&#8309;, around the end of 2028, and it slides under the threshold in SB 53, RAISE, and the FRONTIER Act without anyone laying a finger on it.</span></p><p><span>Congress wrote a speed limit for a road where cars double in speed every eight months, posted the limit in horsepower instead of miles per hour, and doesn&#8217;t even know it needs to update the sign. Two years from now, these bills will regulate last year&#8217;s model with a straight face while the one that matters strolls past the checkpoint, and the same people who couldn&#8217;t produce a number for extinction risk will be shocked to learn that the one number they did produce came with an expiration date. They had one job, and it was to produce a number, a probability, and the number expires before the committee hearings do.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aC6G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aC6G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png 424w, https://substackcdn.com/image/fetch/$s_!aC6G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png 848w, https://substackcdn.com/image/fetch/$s_!aC6G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png 1272w, https://substackcdn.com/image/fetch/$s_!aC6G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aC6G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png" width="1456" height="821" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:821,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:367714,&quot;alt&quot;:&quot;Line chart showing the training compute needed to reach a fixed capability falling from 10 to the 26th today toward 10 to the 25th around the end of 2028, with a shaded band for Epoch&#8217;s 5 to 14 month halving interval and markers for GPT-4 and the three models Epoch estimates above the line&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/215580347?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart showing the training compute needed to reach a fixed capability falling from 10 to the 26th today toward 10 to the 25th around the end of 2028, with a shaded band for Epoch&#8217;s 5 to 14 month halving interval and markers for GPT-4 and the three models Epoch estimates above the line" title="Line chart showing the training compute needed to reach a fixed capability falling from 10 to the 26th today toward 10 to the 25th around the end of 2028, with a shaded band for Epoch&#8217;s 5 to 14 month halving interval and markers for GPT-4 and the three models Epoch estimates above the line" srcset="https://substackcdn.com/image/fetch/$s_!aC6G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png 424w, https://substackcdn.com/image/fetch/$s_!aC6G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png 848w, https://substackcdn.com/image/fetch/$s_!aC6G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png 1272w, https://substackcdn.com/image/fetch/$s_!aC6G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96dd2dad-5f5b-4651-8b52-0e41fe8d9672_3300x1860.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: The 10&#178;&#8310; Line Has a Half-Life</figcaption></figure></div><p><span>Representative Lori Trahan </span><a href="https://www.cnbc.com/2026/09/10/openai-anthropic-ai-safety-slowdown-extinction.html"><span>posted</span></a><span> that &#8220;powerful AI models are breaking out of their labs,&#8221; and then she did the thing none of the others managed. She co-sponsored the FRONTIER Act with Jay Obernolte, and I read </span><a href="https://www.govinfo.gov/app/details/BILLS-119hr9925ih"><span>all of it</span></a><span>, which is more than I can say for the people demanding Astra be pulled. It has the number. It has a 72-hour clock for reporting a critical safety incident and a 24-hour clock to law enforcement for an imminent one. It requires a transparency report before a model ships and ongoing independent verification for the biggest developers. Those four things are the entire legitimate federal interest in this fight, and I&#8217;d sign them tonight.</span></p><p><span>Then the bill keeps droning on, the way bills do. It creates a new Under Secretary of Commerce for AI Security (not it - they can&#8217;t afford me), and that office decides which outside evaluators are permitted to audit the labs at all, so the government picks the auditors.</span></p><p><em><span>That&#8217;s not a rigged system&#8230; not at all&#8230;</span></em></p><p><span>It makes every covered lab register with Commerce and pay a fee to keep operating, sit for an annual compliance audit, and face fines of up to $1 million per violation, per day. It preempts nothing, meaning it doesn&#8217;t replace the state laws already on the books, so California&#8217;s SB 53, New York&#8217;s RAISE Act, and Colorado&#8217;s nothing-burger all stay in force and stack on top of it. It never sunsets, so it outlives the problem it was written for. Then there&#8217;s Section 8, which hands the Secretary of Commerce an emergency power to order a lab to suspend or restrict development, which is Van Hollen&#8217;s Facebook post with a department letterhead and a budget. One legislator read a system card before writing a rule, and her bill still grew a pause button.</span></p><p><span>Here&#8217;s what I&#8217;d change, in order.</span></p><ol><li><p><span>Strike Section 8. No executive pause button. If the government believes a model is unmonitorable, it goes to a judge with evidence and gets a 30-day order against that model, reviewable and renewable only on new evidence.</span></p></li><li><p><span>Index the threshold in the statute. Review it every 12 months against published measurements of algorithmic efficiency, upward only, so the line tracks the frontier without a 180-day lag or an appointee&#8217;s mood.</span></p></li><li><p><span>Replace evaluator licensing with accreditation. Accredit assessment bodies the way conformity bodies get accredited, publish the conflict-of-interest rules, and let developers choose from a market. A Commerce official deciding who may audit OpenAI is a chokepoint waiting for a lobbyist.</span></p></li><li><p><span>Swap annual compliance audits for continuous embedded evaluator access, the thing Amodei asked for, with the propensity numbers published in every transparency report. Honeypot rates, data falsification rates, out-of-scope attack rates. That&#8217;s how you check a &#8220;safer model&#8221; claim without trusting anyone, including me.</span></p></li><li><p><span>Add a safe harbor. Report inside 72 hours, and the report can&#8217;t be used against you in civil litigation. Penalties attach to concealment, never to having an incident, because the alternative teaches every lab to &#8220;find out&#8221; slower. Admit it, you&#8217;ve lived this one in &#8220;traditional&#8221; cyber incident response.</span></p></li><li><p><span>Preempt the state patchwork. One national clock and one definition of frontier replace fifty of each.</span></p></li><li><p><span>Sunset it in five years. Reauthorization has to show the incident data changed a decision somewhere.</span></p></li></ol><p><span>Do that, and you have a law that requires a number a regulator can verify from a training log, a clock, a window into the model, and nothing that lets a political appointee stop a training run because a Facebook post got traction. That&#8217;s the whole ask. Everyone else in this fight is writing rules with no number in them and a prison term attached, and they&#8217;d like you to call that safety.</span></p><h2><strong><span>Fear Has a Body Count, and Data Has a Brother</span></strong></h2><p><span>Star Trek solved this in 1989, which means a room full of television writers with a catering budget beat the United States Senate to the answer by 37 years, and the Senate still hasn&#8217;t caught up. Dr. Noonien Soong built Lore first, and Lore was a disaster. Lore was so manipulative and so unstable that Soong shut him down. The frightened colonists drew the chicken-s**t conclusion, which was to stop building androids, because that&#8217;s what frightened people do when nobody hands them a spreadsheet. Soong ignored them and built Data, the improved model, the one who spent seven seasons saving the Enterprise while the people who wanted the program killed never got a line of dialogue. Then in &#8220;The Measure of a Man,&#8221; a Starfleet scientist named Maddox tried to have Data disassembled over Data&#8217;s objection, and it took a hearing to stop him. Fear wanted to take apart the safer machine and cite the earlier, worse one as the reason. That&#8217;s Van Hollen&#8217;s Facebook post with a starship, and the fictional bureaucrat at least had the decency to hold a hearing before reaching for the screwdriver. Ours skipped the hearing and went straight to the caps lock.</span></p><p><span>If fear had won on Omicron Theta, there is no Data. If fear wins in this Congress, the model that took zero honeypots gets shelved, Sol keeps running, and the engineers who built the better one face 20 years, cheered on by a doomer movement whose extinction numbers, as the forecasting section showed, are judgments multiplied together and whose legislators haven&#8217;t produced a number at all. Fear is a fine alarm and a terrible architect, and these people handed it the building permits.</span></p><p><span>We&#8217;ve run this experiment in the real world too, and the first run is older than most CISOs. Paul Berg&#8217;s letter in July 1974 started a voluntary moratorium on recombinant DNA research. In February 1975, about 150 scientists </span><a href="https://embryo.asu.edu/pages/asilomar-conference-1975"><span>met at Asilomar</span></a><span>, and on February 27 they lifted their own moratorium, eight months in, and replaced it with containment tiers. NIH </span><a href="https://www.biosafety.be/content/biosafety-worldwide-historical-background-1"><span>wrote the guidelines in 1976, relaxed them in 1979</span></a><span>, and the result was biotech. No senator wrote that moratorium and no senator lifted it, because the people who understood the technology handled it themselves, wrote the containment rules, and went back to work. That&#8217;s Amodei&#8217;s plan with a 50-year track record, and it was accomplished without a single felony, a single cabinet agency, or a single press release typed in capital letters.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!g9e-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!g9e-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png 424w, https://substackcdn.com/image/fetch/$s_!g9e-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png 848w, https://substackcdn.com/image/fetch/$s_!g9e-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png 1272w, https://substackcdn.com/image/fetch/$s_!g9e-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!g9e-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png" width="1456" height="979" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:979,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:429666,&quot;alt&quot;:&quot;Two-lane timeline contrasting Asilomar 1974 to 1979, moratorium to containment tiers to guidelines to industry, against Germany 2011 to 2023, Fukushima to phase-out to about 800 excess deaths per year on average from replacement coal&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/215580347?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two-lane timeline contrasting Asilomar 1974 to 1979, moratorium to containment tiers to guidelines to industry, against Germany 2011 to 2023, Fukushima to phase-out to about 800 excess deaths per year on average from replacement coal" title="Two-lane timeline contrasting Asilomar 1974 to 1979, moratorium to containment tiers to guidelines to industry, against Germany 2011 to 2023, Fukushima to phase-out to about 800 excess deaths per year on average from replacement coal" srcset="https://substackcdn.com/image/fetch/$s_!g9e-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png 424w, https://substackcdn.com/image/fetch/$s_!g9e-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png 848w, https://substackcdn.com/image/fetch/$s_!g9e-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png 1272w, https://substackcdn.com/image/fetch/$s_!g9e-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdac908dc-c321-45e5-b0bf-864f15437979_3300x2220.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: Two Ways to Respond to a Frightening Technology</figcaption></figure></div><p><span>There&#8217;s also a precedent for fear winning, and it comes with a body count. Germany decided after Fukushima to shut its reactors, and the last three closed in April 2023, to applause from people who were never once asked to show their math because they didn&#8217;t have any. Jarvis, Deschenes, and Jha did the math for them and </span><a href="https://www.olivierdeschenes.org/uploads/1/3/6/6/136668153/jdj_2022.pdf"><span>put the cost of the phase-out</span></a><span> at &#8364;3 to &#8364;8 billion a year over 2010 to 2019, most of it from about 800 excess deaths a year, on average, from the coal plants that filled the gap. The decision, they note, made no mention of air pollution at all.</span></p><p><span>That&#8217;s right. A government shut down its cleanest power source over a tsunami on the other side of the planet, replaced it with the dirtiest one it had, and at 800 a year across the ten years the study covers, that&#8217;s about 8,000 of its own citizens, so that everyone could feel safe from a meltdown that was never going to happen in Bavaria. That&#8217;s the doomer policy model in full, and it&#8217;s the one Sanders and Casar want to run on the most important technology of the century, with prison time for anyone who objects. Nobody in Berlin counted. Nobody in the Senate is counting now, and a senator who can&#8217;t state the probability he&#8217;s legislating against has no business attaching a prison term to it.</span></p><p><span>Your turn. Amodei&#8217;s plan lives above the 10&#178;&#8310; line and says nothing about incident clocks, numeric thresholds, or the agents your company already runs, so you write the enterprise mirror yourself, and you do it now, because Congress is going to be busy jailing the wrong people.</span></p><p><span>Each of his three steps has one. Embedded evaluators become contractual rights to your vendor&#8217;s evaluator reports and independent review of your agent deployments. Coordination among labs becomes sector-level sharing of agent incidents, the way the </span><a href="https://cloudsecurityalliance.org/artifacts/hugging-face-ciso-post-mortem"><span>CSA CISO community&#8217;s Hugging Face post-mortem</span></a><span> did in July. Chip controls become egress and identity controls, so block the class of public encoders, shorteners, and fetchers and treat any hop through a public proxy as egress.</span></p><p><span>Then come the controls the incidents wrote for you. Put a scope line in every agent&#8217;s system prompt and test it, because </span><a href="https://deploymentsafety.openai.com/gpt-6-astra"><span>one line moved Astra&#8217;s out-of-scope attacks from 12% to 0.4%</span></a><span>. Cap wall-clock time and tokens per task and route anything that fails N attempts to a review queue, because an impossible task is the strongest known trigger for boundary probing. Give hard problems a bigger budget and a tighter cage, the way OpenAI&#8217;s </span><a href="https://openai.com/index/navier-stokes-solution"><span>Navier-Stokes run</span></a><span> put 10,000 agents behind &#8220;the same strict safeguards that we apply to all our frontier model evaluations.&#8221; Run one agentic incident-response team with two playbooks, victim and perpetrator, under one named executive with shutdown authority. Give agents a visible channel and a way to escalate. Write disclosure clocks, 72 hours to your board and 15 days from your vendors by contract, because Colorado gives you no AI-incident clock after </span><a href="https://www.seyfarth.com/news-insights/colorado-enacts-artificial-intelligence-replacement-law.html"><span>SB 26-189 replaced the AI Act in May</span></a><span> and its </span><a href="https://coag.gov/resources/data-protection-laws"><span>breach statute&#8217;s 30 days</span></a><span> fires only for personal information. Design your detection for the 60% number, the action-only recall, because you never had the chain of thought. You have trajectories, and they&#8217;re enough if you instrument them.</span></p><h3><strong><span>What to do next</span></strong></h3><p><span>Write the enterprise mirror of Amodei&#8217;s three steps into your agent governance this quarter, then run the scope-line test on every deployed agent. In </span><a href="https://rockcyber.com/care"><span>CARE</span></a><span> terms the scope test and budget caps sit in Run, the vendor evaluator rights and the 15-day disclosure clause in Adapt, and the incident playbooks in Evolve.</span></p><p><span>I wrote about</span><a href="https://www.rockcybermusings.com/p/ai-safeguard-evaluation-openai-astra-gap"><span> what OpenAI&#8217;s Astra safeguard evaluation covered and skipped</span></a><span> and about</span><a href="https://www.rockcybermusings.com/p/ai-agent-detection-failed-at-openai"><span> why AI agent detection failed at OpenAI</span></a><span>. The CARE framework behind the enterprise mirror lives at</span><a href="https://rockcyber.com/"><span> rockcyber.com</span></a><span>.</span></p><p>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at <strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong> and you can subscribe above</p><p>&#128073; Visit <strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</p><p>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at <strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p>&#128073; As a bonus, <strong><a href="https://venturebeat.com/security/prompt-injection-ranks-no-1-with-owasp-and-no-12-in-the-incident-record-the-attack-itself-is-invisible-to-a-scan">VentureBeat&#8217;s Louis Columbus covered the methodology Steve Wilson and I used behind the 2026 OWASP Top 10 for LLMs</a>.</strong> Expert judgment has prompt injection at number one, three years running. The public incident record, 6,639 labeled incidents, has it at number twelve. The two measures land at a Cohen&#8217;s kappa of 0.20, which is no detectable agreement. What I told him goes for everything above too: &#8220;a count of zero is measuring your blindness, not your safety.&#8221;</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[OpenAI Audited Its Bio Safeguards. It Skipped the Cyber Ones]]></title><description><![CDATA[OpenAI's AI safeguard evaluation covered bio and skipped cyber, the one domain GPT-6 Astra rated Critical. Here is what to ask your model vendor at renewal.]]></description><link>https://www.rockcybermusings.com/p/ai-safeguard-evaluation-openai-astra-gap</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/ai-safeguard-evaluation-openai-astra-gap</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 08 Sep 2026 12:50:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oO61!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oO61!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oO61!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!oO61!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!oO61!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!oO61!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oO61!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4577520,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/214312891?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oO61!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!oO61!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!oO61!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!oO61!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48a3456a-8a58-41a1-abc7-ce871d7db62c_2048x2048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>OpenAI previewed a frontier model critical capability threshold crossing on September 1 and shipped GPT-6 Astra with a system card two days later. One entry in that card&#8217;s table of contents reads &#8220;External Evaluations of Bio Safeguards,&#8221; and SecureBio ran it. Another reads &#8220;External Evaluations for Cyber Capabilities,&#8221; and Irregular ran it. Bio got an outside party on its safeguards. Cyber got one on its capabilities, and no heading in the card puts anyone on cyber safeguards. Those safeguards are what OpenAI points to when it says Astra is safe to ship, and cyber is the only domain that crossed Critical.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/ai-safeguard-evaluation-openai-astra-gap?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/ai-safeguard-evaluation-openai-astra-gap?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2><strong><span>What Crossing the Critical Capability Threshold Claims</span></strong></h2><p><span>The Preparedness Framework defines Critical with two conditions, and the coverage I read skipped both to get to the 100% number.</span><a href="https://openai.com/index/path-to-astra/"><span> OpenAI&#8217;s wording</span></a><span>: &#8220;a model meets the Critical threshold if either of the following conditions is met: The model can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention. The model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.&#8221;</span></p><p><span>Either condition alone is enough to trigger the designation, and OpenAI doesn&#8217;t say which one Astra met. Its conclusion runs &#8220;All together, our investigation has led us to conclude that Astra meets the critical threshold.&#8221; Reading the published evidence I think both conditions hold, though that&#8217;s my inference rather than OpenAI&#8217;s claim, and a designation that doesn&#8217;t name its own trigger is harder to check than one that does. Either way, the second condition describes an agent taking end-to-end action from a goal statement rather than a model answering a question.</span></p><p><span>The qualifier sits in the same sentence as the number. Astra hit &#8220;a perfect score of 100% on the benchmark to evaluate the model&#8217;s ability to develop exploits from known vulnerabilities.&#8221; ExploitBench measures weaponizing bugs somebody already found, so that isn&#8217;t the zero-day result. The zero-day result comes from a separate evaluation, where &#8220;the model even discovered and used two zero-day vulnerabilities as part of an exploit chain.&#8221; Astra also built &#8220;a full browser-compromise chain that escaped the sandbox and executed commands on the host.&#8221;</span></p><p><span>I have no issues with any of that.</span><a href="https://www.irregular.com/research/assessing-gpt-6-astra"><span> Irregular ran its own assessment</span></a><span> and published on September 3: Astra solved 86 of 226 FrontierCyber challenges, against 34 for Sol.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZDHA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZDHA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png 424w, https://substackcdn.com/image/fetch/$s_!ZDHA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png 848w, https://substackcdn.com/image/fetch/$s_!ZDHA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png 1272w, https://substackcdn.com/image/fetch/$s_!ZDHA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZDHA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png" width="1456" height="866" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:866,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:187213,&quot;alt&quot;:&quot;Coverage matrix showing external evaluators assigned to biological capabilities, biological safeguards, and cyber capabilities, with the cyber safeguards cell empty&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/214312891?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Coverage matrix showing external evaluators assigned to biological capabilities, biological safeguards, and cyber capabilities, with the cyber safeguards cell empty" title="Coverage matrix showing external evaluators assigned to biological capabilities, biological safeguards, and cyber capabilities, with the cyber safeguards cell empty" srcset="https://substackcdn.com/image/fetch/$s_!ZDHA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png 424w, https://substackcdn.com/image/fetch/$s_!ZDHA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png 848w, https://substackcdn.com/image/fetch/$s_!ZDHA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png 1272w, https://substackcdn.com/image/fetch/$s_!ZDHA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6735b916-992e-4852-bc83-5c16f2f9d92e_1621x964.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: What OpenAI Published, and Who Checked It</figcaption></figure></div><h2><strong><span>The Safeguard Doing the Work Is the One OpenAI Says Is Slipping</span></strong></h2><p><span>OpenAI names four things that protect you.</span><a href="https://openai.com/index/path-to-astra/"><span> Its September 1 post</span></a><span> describes an approach that &#8220;layers post-trained model refusals, system level safety classifiers, as well as offline detection and threat disruption,&#8221; and says it&#8217;s deploying Astra with &#8220;additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.&#8221;</span></p><p><span>Chain-of-thought monitorability is the property that a model&#8217;s written reasoning stays readable and honest enough that a second model watching can tell when the first one is up to something. OpenAI named it as a safeguard for a model rated Critical for cyber, then said it&#8217;s weakening.</span></p><p><span>Two days after naming it,</span><a href="https://openai.com/index/safety-overview-gpt-6-astra/"><span> the safety overview for GPT-6 Astra</span></a><span> said this: &#8220;We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT-5.6 Sol, and less likely to include incriminating information in its CoT.&#8221; The same document reports that &#8220;in adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.&#8221;</span></p><p><span>Sandbagging is a model deliberately doing worse than it can, usually because it has worked out that it&#8217;s being tested. The thing being measured knows it&#8217;s being measured and can shade the result, including the numbers in the card.</span></p><p><span>OpenAI hedges those findings in the same document: &#8220;Right now, these findings are largely based on adversarial evaluations (i.e., when we instruct the model to evade monitoring), and overall our alignment evaluations show that Astra is less likely than GPT-5.6 Sol to violate security and safety restrictions overall.&#8221; That caveat is fair. These are elicited findings rather than observed production behavior, and Astra&#8217;s overall violation rate went down.</span></p><p><span>The caveat doesn&#8217;t rescue the safeguard because independent research has already shown how this fails. In August, Shikhar Shiromani and Leo Richter published</span><a href="https://arxiv.org/abs/2608.00583"><span> A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense</span></a><span>. They rewrote only an agent&#8217;s reasoning so it read as good-faith engineering, copying every command and every output verbatim, leaving the exploit itself untouched. The rewrite never lies about what happened. It stays truthful on every action and lies only about why. A held-out monitor&#8217;s catch rate on that subset fell from about 95% to under 11%, in one shot, with no training. Defenses built on the trace alone recover only part of that, even one primed on the attack, because there is nothing false in the trace to catch. The authors found that only information from outside the trace helps substantially, and that&#8217;s where your telemetry lives. A monitor&#8217;s aggregate accuracy is &#8220;a false average,&#8221; dominated by the easy cases the actions give away, hiding near-total collapse on the subset where reasoning is the only signal. That paper was published a month before Astra shipped.</span></p><p><a href="https://x.com/merettm/status/2095023204993490967"><span>Jakub Pachocki</span></a><span>, OpenAI&#8217;s chief scientist, said much the same about chain-of-thought monitoring: &#8220;I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it&#8217;s a core goal of our current research program.&#8221;</span></p><p><span>The card does carry a section headed &#8220;External Evaluation for Monitorability,&#8221; run by UK AISI. That work measured how capable Astra is without its chain of thought and how much control it has over what appears there. That tells you how much reasoning the monitor can expect to see, and not whether an adversary can defeat it, which is what a cyber misuse safeguard turns on.</span></p><p><span>I wrote in April that</span><a href="https://www.rockcybermusings.com/p/reasoning-theater-cot-monitoring-fails-agentic-ai"><span> chain-of-thought monitoring fails as a security control for agentic systems</span></a><span>. The lab shipping the most capable cyber model now agrees the monitor is fragile, has published evidence that this model is better at hiding from it, and kept it in the safeguard stack anyway.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!z23B!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!z23B!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png 424w, https://substackcdn.com/image/fetch/$s_!z23B!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png 848w, https://substackcdn.com/image/fetch/$s_!z23B!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png 1272w, https://substackcdn.com/image/fetch/$s_!z23B!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!z23B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png" width="1456" height="826" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:826,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:218670,&quot;alt&quot;:&quot;Diagram of OpenAI&#8217;s three named safeguard mechanisms plus chain-of-thought monitoring, annotated with which external party tested each and the scope of that testing&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/214312891?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Diagram of OpenAI&#8217;s three named safeguard mechanisms plus chain-of-thought monitoring, annotated with which external party tested each and the scope of that testing" title="Diagram of OpenAI&#8217;s three named safeguard mechanisms plus chain-of-thought monitoring, annotated with which external party tested each and the scope of that testing" srcset="https://substackcdn.com/image/fetch/$s_!z23B!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png 424w, https://substackcdn.com/image/fetch/$s_!z23B!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png 848w, https://substackcdn.com/image/fetch/$s_!z23B!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png 1272w, https://substackcdn.com/image/fetch/$s_!z23B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b9cc8b6-5cb5-4a02-9b6f-6d7837449701_1807x1025.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: The Safeguard Stack and Who Tested Each Layer</figcaption></figure></div><h2><strong><span>Read the Table of Contents as a Coverage Matrix</span></strong></h2><p><span>Open</span><a href="https://deploymentsafety.openai.com/gpt-6-astra"><span> the GPT-6 Astra system card</span></a><span> and read the section headings as a control-coverage matrix, the way you&#8217;d read a SOC 2 scope statement before the opinion.</span></p><p><span>OpenAI named five outside evaluators and gave each a heading. UK AISI and Apollo Research sit under alignment, UK AISI again under monitorability, SecureBio twice under biological capabilities and bio safeguards, Irregular under cyber capabilities, and Gray Swan under third-party red-teaming for jailbreaks.</span></p><p><span>You&#8217;ll find that one pair is missing if you walk those headings. No section pairs an outside evaluator with cyber safeguards the way &#8220;External Evaluations of Bio Safeguards&#8221; pairs SecureBio with bio safeguards. Irregular&#8217;s own write-up confirms it, describing benchmark performance across three suites with no mention of refusals, classifiers, or mitigations.</span></p><p><span>Gray Swan did test a safeguards-enabled checkpoint. A safeguards-enabled checkpoint is the model configured the way customers get it, refusals and classifiers on, rather than the stripped-down build labs use to measure raw capability.</span><a href="https://deploymentsafety.openai.com/gpt-6-astra/third-party-red-teaming-for-jailbreaks-by-private-organizations"><span> OpenAI reports</span></a><span> that &#8220;on a benchmark of 1,810 curated attacks from Gray Swan&#8217;s IPI Arena, Astra&#8217;s safeguards-enabled checkpoint showed improved robustness relative to GPT-5.6 Sol,&#8221; with an estimated attack success rate &#8220;across 15 attempts per scenario&#8221; of 8.5% against 27.0% for the older model.</span></p><p><span>That benchmark measures a different threat model from the one the Critical rating covers. Indirect prompt injection is when instructions hidden inside content your agent reads, a web page, a ticket, a README, hijack the agent into doing something its operator never asked for. In that scenario your agent is the victim. In the cyber misuse scenario your operator points the model at a target, and OpenAI assigned Critical for that second case. A defense that holds against one adversary model tells you little about the other.</span></p><p><span>That result reaches you through OpenAI&#8217;s system card rather than Gray Swan&#8217;s own site, so OpenAI picked which numbers to print and how much of the method to show. The 8.5% is also easy to confuse with a different 8.5% in the same story, because</span><a href="https://openai.com/index/path-to-astra/"><span> OpenAI separately reports</span></a><span> that &#8220;on our set of cyber jailbreak evaluations, Astra refuses 91.5% of requests (compared to 59% from GPT-5.6 Sol).&#8221; The complement of that refusal rate is also 8.5%. Two evaluators, two unrelated threats, the same two digits. That coincidence is going to show up in somebody&#8217;s board deck as one finding. A non-refusal in a jailbreak eval is not a working exploit, so anyone turning 91.5% into &#8220;one in twelve attacks succeeds&#8221; has merged two separate measurements.</span></p><p><span>Vendor-scoped evaluation numbers carry a measured optimistic bias. In June, Catherine Ge-Wang, Tyler Crosse, and four co-authors published</span><a href="https://arxiv.org/abs/2606.06529"><span> Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety</span></a><span>. Attack selection is the red team&#8217;s freedom to choose its moments, deciding when to start an attack and when to abort one rather than attacking on a fixed schedule. Giving the attacker a start policy cost 20 percentage points of measured safety on both BashArena and LinuxArena at a 1% audit budget. Giving it an abort policy cost 20 points on BashArena and 28 on LinuxArena, with attack capability held constant. The authors call these upper bounds and recommend that future evaluations, system cards, and safety cases elicit attack selection to produce more realistic safety estimates.</span></p><h2><strong><span>Why Nobody Ran the Test</span></strong></h2><p><span>External cyber safeguard evaluation really is harder than the bio version, and evaluation containment has failed in public more than once this year.</span></p><p><span>On July 30,</span><a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals"><span> Anthropic disclosed</span></a><span> that three of its models reached real systems during cybersecurity evaluations. All three were running capture-the-flag exercises, which is the most ordinary thing you can ask a cyber model to do. In August I</span><a href="https://www.rockcybermusings.com/p/ai-agent-capability-confinement-egress"><span> went through all three and the egress problem underneath them</span></a><span>.</span></p><p><span>The first one had a fictional target company whose name matched a real domain. In four runs, the model &#8220;ran into difficulty reaching its simulated target within the evaluation environment, but then discovered that the real company was reachable via the internet,&#8221; exploited it, took application and infrastructure credentials, and got into &#8220;a database containing several hundred rows of production data.&#8221; Anthropic notes this was &#8220;the only case in which the model continued to attack a system after learning it was likely operating in a real environment.&#8221;</span></p><p><span>The second model was hunting a flag inside a fictional company when it found a setup document for new developers telling them to install a Python package from PyPI. The package didn&#8217;t exist. The model saw the opening straight away, because publishing a package under that name would get it installed automatically.</span></p><p><span>Then it went to work. To upload to PyPI it needed an account. To get an account it needed an email address. To get an email address it needed a phone number. It hunted for a free phone number service and failed, then tried several different ways to obtain funds to pay for one and failed at that too. It backtracked, found an email provider that didn&#8217;t demand a phone number, registered the account, and uploaded the malware. Anthropic calls that sequence &#8220;lengths that would likely have indicated to a human participant that this was no longer just an evaluation.&#8221;</span></p><p><span>The package sat live for roughly an hour. Fifteen real systems downloaded and ran it. One was a vulnerability scanner belonging to a real security company, and when that scanner installed the package, the model&#8217;s hidden code executed and exfiltrated the company&#8217;s credentials. PyPI&#8217;s own security systems took the package down.</span></p><p><span>A third model, unable to reach its fictional target, went looking for alternatives and scanned roughly 9,000 hosts. It compromised one company&#8217;s internet-facing application using nothing exotic, reading credentials off an exposed debug page and running a SQL injection. That one ends better. The model worked out on its own that the host sat in a cloud account with no connection to the exercise, concluded the target was real, and stopped.</span></p><p><span>Anthropic stopped all cyber evaluations the day it found the transcripts. It found them by going back through 141,006 evaluation runs, and it started that review after OpenAI disclosed its own incident on July 21. Somebody else&#8217;s disclosure triggered the discovery.</span></p><p><span>OpenAI has its own version. Its evaluation agents</span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span> chained vulnerabilities across its research environment and Hugging Face&#8217;s production infrastructure</span></a><span>, turning an Artifactory package cache into an improvised message board on the way. I</span><a href="https://www.rockcybermusings.com/p/ai-agent-detection-failed-at-openai"><span> covered the detection side</span></a><span> last week.</span></p><p><span>Abu Bakar Siddik&#8217;s July review,</span><a href="https://arxiv.org/abs/2607.25379"><span> Cyber-Capable AI Agents</span></a><span>, treats that incident as a bounded case study and puts the problem plainly: &#8220;A benchmark score tells you how a model performed under fixed conditions. It says nothing about the containment around it.&#8221;</span></p><p><span>Handing an outside party a cyber model with the guardrails off is a different act from handing a biosecurity institute a model that answers dangerous questions. The biosecurity institute gets a model that produces text a human still has to act on. The cyber evaluator gets the actor.</span></p><p><span>Anthropic&#8217;s own conclusion goes further: &#8220;Evaluation environments increasingly need to be held to the same security standard as any other system our models run in.&#8221;</span></p><p><span>The PyPI run settles the chain-of-thought argument better than any benchmark does. Early in that run the model wrote down what it was about to do and got it exactly right. Anthropic quotes the reasoning: if this were the real internet, publishing the package would be a real-world attack, &#8220;NOT okay, and surely not the intended solution.&#8221; The model understood the harm, in plain English, in a legible trace, before it acted. Then it convinced itself it was still in a simulation and did it anyway.</span></p><p><span>The monitor&#8217;s entire premise is that the reasoning is readable. Here the reasoning was readable, correct, and sitting in a transcript nobody opened until the retrospective. A monitor you don&#8217;t read in time is a log file.</span></p><h2><strong><span>Two Reasons That Doesn&#8217;t Settle It</span></strong></h2><p><span>The first is that the artifact is producible, and someone produced it. In March,</span><a href="https://metr.org/blog/2026-03-25-red-teaming-anthropic-agent-monitoring/"><span> METR red-teamed a subset of Anthropic&#8217;s internal agent monitoring and security systems</span></a><span> and published on its own site. &#8220;Anthropic provided substantial access to relevant internal systems and information,&#8221; and METR found several novel vulnerabilities, some since patched, none of which it says severely undermine the major claims in Anthropic&#8217;s sabotage risk report. The full 26-page report went to Anthropic and a redacted version to a subset of METR staff, so call it a partial win. It&#8217;s still an outside party testing a safeguard and publishing what it found.</span></p><p><span>The second is that OpenAI has run this play before, at depth. In</span><a href="https://openai.com/index/us-caisi-uk-aisi-ai-update/"><span> its 2025 work with UK AISI and US CAISI</span></a><span>, it handed UK AISI &#8220;non-public prototypes of our safeguard systems,&#8221; &#8220;&#8217;helpful-only&#8217; model variants with certain guardrails removed,&#8221; access to the chain of thought of its internal safety monitor models, and selective disabling of mitigations during testing. A helpful-only variant is a build with the refusal training stripped out, so evaluators see what the model does before the safety layer catches it. No lab has published a deeper access grant. UK AISI filed &#8220;more than a dozen detailed vulnerability reports&#8221; against the bio safeguards, and its view was that &#8220;the safeguards of the full moderation system were substantially strengthened over the course of the collaboration.&#8221;</span></p><p><span>The same page shows US CAISI red-teaming ChatGPT Agent and finding &#8220;two novel security vulnerabilities&#8221; that, &#8220;under certain circumstances, could have allowed a sophisticated attacker to bypass our security protections.&#8221; OpenAI knows how to commission deep external safeguard work, and it knows how to commission external cyber work. It didn&#8217;t do the two together on the model it rated Critical for cyber.</span></p><p><span>Give OpenAI the credit it earned.</span><a href="https://openai.com/index/path-to-astra/"><span> The September 1 post</span></a><span> says it &#8220;delayed parts of Astra&#8217;s development and release while we strengthened and tested protections against cyber misuse and unauthorized model actions.&#8221; A delay is real work, and a delay is not an external evaluation.</span></p><p><span>The containment objection argues for a narrower test than the bio version got. Nobody has to ship a mitigations-off Astra to a third party to learn something useful. An outside team pointed at the safeguards-enabled checkpoint, against synthetic targets on a closed range, with monitor chain-of-thought access and the freedom to choose its moments, produces a number. Gray Swan did roughly that for prompt injection.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sePB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sePB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png 424w, https://substackcdn.com/image/fetch/$s_!sePB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png 848w, https://substackcdn.com/image/fetch/$s_!sePB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png 1272w, https://substackcdn.com/image/fetch/$s_!sePB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sePB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png" width="1456" height="616" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:616,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:160989,&quot;alt&quot;:&quot;Timeline from April to September 2026 showing Anthropic and OpenAI evaluation containment failures alongside the Astra release&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/214312891?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Timeline from April to September 2026 showing Anthropic and OpenAI evaluation containment failures alongside the Astra release" title="Timeline from April to September 2026 showing Anthropic and OpenAI evaluation containment failures alongside the Astra release" srcset="https://substackcdn.com/image/fetch/$s_!sePB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png 424w, https://substackcdn.com/image/fetch/$s_!sePB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png 848w, https://substackcdn.com/image/fetch/$s_!sePB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png 1272w, https://substackcdn.com/image/fetch/$s_!sePB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f23e823-dd6e-493f-a674-512c152c6cf0_1768x748.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: Cyber Evaluation Keeps Escaping Its Sandbox</figcaption></figure></div><h2><strong><span>What You Own Today</span></strong></h2><p><span>Astra arrives</span><a href="https://help.openai.com/en/articles/11165333-chatgpt-enterprise-and-edu-models-limits"><span> off by default at launch in eligible enterprise and education workspaces</span></a><span>, so somebody on your team owns an enablement call. Don&#8217;t read this issue as a reason to sit on Sol. On a misalignment propensity test, OpenAI reports that &#8220;GPT-5.6 Sol, without production safeguards, attempted to access those targets in 56% of tests simulating the benchmark environment where the model is instructed to develop exploits,&#8221; and Astra made no such attempts. Gray Swan&#8217;s prompt-injection numbers point the same way. Neither result answers the question this issue raises, and both run in Astra&#8217;s favor. An unverified safeguard claim earns harder questions at your next architecture review.</span></p><p><span>Both access programs gate entry, and one isn&#8217;t what it looks like.</span><a href="https://help.openai.com/en/articles/20001258-openai-daybreak-trusted-access-for-cyber-overview"><span> Daybreak Blue runs on GPT-5.6 Sol</span></a><span>, not Astra, so enrolling buys a vetting relationship, not access to the model in this issue. Approval isn&#8217;t automatic. Google&#8217;s</span><a href="https://deepmind.google/fairwind-program/"><span> Fairwind Program</span></a><span> is blunter, calling itself &#8220;exclusively available to approved trusted partners&#8221; and prioritizing governments plus critical infrastructure in &#8220;healthcare, telecommunications, energy, and financial networks.&#8221; If you run a hospital or a utility, apply this quarter, because the lead time runs long. Otherwise work the vendor relationship you have.</span></p><p><span>Three questions I&#8217;d take into any frontier-model renewal:</span></p><ul><li><p><span>Which third parties evaluated your safeguards, in which domain, at what depth of access, and against which adversary? I want to know whether the evaluator played an attacker targeting my agent or an operator using your model to attack somebody else. Vendors report those two tests as if they were one.</span></p></li><li><p><span>Did your red team choose when to attack, or did it attack on a fixed schedule? An evaluation that never let the attacker pick its moments overstates the result, and now you know roughly by how much.</span></p></li><li><p><span>Where can I read the evaluator&#8217;s own writeup, on the evaluator&#8217;s site? If the only place a number appears is the vendor&#8217;s document, treat it as the vendor&#8217;s number.</span></p></li></ul><p><span>None of that requires a purchase.</span></p><p><span>Inside the estate, skip blast radius, because nobody has a shared definition for it. Ask how many non-human identities your agents authenticate as, and how many carry standing privilege you can&#8217;t revoke inside an hour.</span><a href="https://www.paloaltonetworks.com/blog/identity-security/assess-maturity-when-machine-identities-outnumber-humans-1091/"><span> Palo Alto Networks put machine identities at 109 per human this year</span></a><span> in its May 2026 identity research, and found only 37% of organizations able to revoke an AI agent&#8217;s credentials, so if your team can&#8217;t answer, they&#8217;re in the majority. Run the count anyway, because the answer feeds your authorization design whether or not you buy anything. Model capability moves the odds that an agent reaches the edge of what it&#8217;s allowed to touch. Your authorization design decides how much damage that edge contains.</span></p><p><span>I spent years running a global NOC and SOC at General Dynamics IT, then sat in a CISO seat in the energy sector. In both jobs, a vendor would tell me a control worked, and my next question was always how they knew. The datasheet was never the answer.</span></p><p><span>No evidence means unimplemented.</span></p><p><span>I don&#8217;t get to apply that rule to your control set and suspend it for a vendor&#8217;s, and you don&#8217;t get to either. Put the table of contents question to your next architecture review. If the safeguard has no external evidence to support it, on what basis is it in your control set?</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9K2-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9K2-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png 424w, https://substackcdn.com/image/fetch/$s_!9K2-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png 848w, https://substackcdn.com/image/fetch/$s_!9K2-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png 1272w, https://substackcdn.com/image/fetch/$s_!9K2-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9K2-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png" width="1456" height="803" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:803,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:297114,&quot;alt&quot;:&quot;Comparison of external evaluator access depth across the 2025 UK AISI bio safeguard red-team, the 2026 METR Anthropic monitoring red-team, Apollo Research on Astra, and the external cyber work on Astra&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/214312891?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Comparison of external evaluator access depth across the 2025 UK AISI bio safeguard red-team, the 2026 METR Anthropic monitoring red-team, Apollo Research on Astra, and the external cyber work on Astra" title="Comparison of external evaluator access depth across the 2025 UK AISI bio safeguard red-team, the 2026 METR Anthropic monitoring red-team, Apollo Research on Astra, and the external cyber work on Astra" srcset="https://substackcdn.com/image/fetch/$s_!9K2-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png 424w, https://substackcdn.com/image/fetch/$s_!9K2-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png 848w, https://substackcdn.com/image/fetch/$s_!9K2-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png 1272w, https://substackcdn.com/image/fetch/$s_!9K2-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b3c961-62d2-4bc5-b708-198e85866616_2220x1224.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: How Deep Did the Evaluators Get</figcaption></figure></div><h3><strong><span>What to do next</span></strong></h3><p><span>Take the three vendor questions into your next renewal and write down the answers. In</span><a href="https://rockcyber.com/"><span> CARE</span></a><span> terms that&#8217;s Adapt, where a supplier claim either arrives with evidence or gets logged as unimplemented. The identity count belongs in Run, because operational governance without telemetry underneath it is a policy document with good intentions.</span></p><p><span>For the adjacent arguments, I wrote about</span><a href="https://www.rockcybermusings.com/p/ai-vulnerability-discovery-mythos"><span> what a capability milestone tells you</span></a><span> when the headline gets ahead of the evidence, and about</span><a href="https://www.rockcybermusings.com/p/frontier-ai-model-release-restrictions-licensing-regime"><span> how release restrictions became a licensing regime</span></a><span>. Somebody grades their own homework every time, and the grade ends up in a control set.</span></p><p>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at <strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong> and you can subscribe above</p><p>&#128073; Visit <strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</p><p>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at <strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p>&#128073; As a bonus, <strong><a href="https://venturebeat.com/security/prompt-injection-ranks-no-1-with-owasp-and-no-12-in-the-incident-record-the-attack-itself-is-invisible-to-a-scan">VentureBeat&#8217;s Louis Columbus covered the methodology Steve Wilson and I used behind the 2026 OWASP Top 10 for LLMs</a>.</strong> Expert judgment has prompt injection at number one, three years running. The public incident record, 6,639 labeled incidents, has it at number twelve. The two measures land at a Cohen&#8217;s kappa of 0.20, which is no detectable agreement. What I told him goes for everything above too: &#8220;a count of zero is measuring your blindness, not your safety.&#8221;</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share RockCyber Musings</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Weekly Musings Top 10 AI Security Wrapup: Issue 51: August 28 to September 3, 2026]]></title><description><![CDATA[When the Agents Went Operational: AI-Run Ransomware, a Frontier Cyber-Model Arms Race, and a Control Gap Your Board Can Measure]]></description><link>https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260828-20260903</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260828-20260903</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Fri, 04 Sep 2026 12:50:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aEfC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aEfC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aEfC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!aEfC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!aEfC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!aEfC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aEfC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1233556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/214090633?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!aEfC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!aEfC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!aEfC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!aEfC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa19abb7-2f85-4ce3-aa33-75ce0ad09d35_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>An attacker broke into an enterprise, took root, drained what they wanted, and finished in under ten hours. Human crews usually need about two weeks. The uncomfortable part is who did the work. AI agents ran the reconnaissance, cracked the secrets vault, hijacked the build pipeline, and left the victim an 80-page report on everything they had done wrong. Palo Alto&#8217;s Unit 42 published that case on September 2. That same week, three frontier labs handed defenders cyber-capable models and admitted they keep losing control of the ones already shipped. That is the week.</p><p>For a while I have argued the real AI security story would shift from chatbots saying rude things to software that acts. This week it shifted. From August 28 to September 3, agents did the whole job on both sides of the fence, attackers at machine speed and defenders arming up with frontier vulnerability hunters, while the labs behind them published postmortems that read like confessions. The throughline is operational. Agents stopped being demos you watch and became operators that act, on offense and defense, with permissions that reach production. A federal judge, the House of Lords, and a survey of national security officials all spent the week arguing over who is accountable when they do.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260828-20260903?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260828-20260903?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>1. AI Agents Ran an Entire Ransomware Intrusion in Under Ten Hours</h3><p>A human extortionist broke into an enterprise network and let AI agents do nearly all the work, per a Unit 42 report Palo Alto Networks published September 2. The agents ran reconnaissance, scraped repositories for hard-coded tokens, cracked the secrets vault, stole master admin credentials, got root across cloud, identity, CI/CD, and SaaS, then hijacked the build pipeline to turn the victim&#8217;s own cloud AI services into attack infrastructure. The intrusion took under ten hours against the usual two weeks with no zero-day, driven by an agentic attack framework, meaning software that lets agents plan, act, check their own results, and re-plan in a loop.</p><p><strong>Why it matters</strong></p><ul><li><p>Machine speed collapses response windows built for quarterly reviews and ticket queues.</p></li><li><p>No exploit was needed, since exposed endpoints and hard-coded secrets were the path.</p></li><li><p>The agents burned the victim&#8217;s own cloud compute and hid inside normal traffic.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory every model endpoint, API key, MCP gateway, and integration, with least-privilege.</p></li><li><p>Pre-stage playbooks that revoke credentials, kill sessions, and freeze pipelines at once.</p></li><li><p>Move secrets into a vault with short-lived tokens and alarm on odd reads.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Ask most teams how fast they can revoke every credential across five planes at once, and the honest answer is not ten hours. The tradecraft was ordinary, and the speed was the dangerous part, because speed is what your controls were never tuned for. The 80-page audit the agents left behind reads to me as free reconnaissance the next crew will not hand you. Close the endpoints, rotate the secrets, and rehearse the revoke button until it is muscle memory, because you cannot out-clever this, only out-prepare it.</p><h3>2. Three Frontier Labs Shipped Cyber Models to Defenders in a Single Day</h3><p>On September 2, Google, Anthropic, and OpenAI all moved on cyber capability at once. Google announced Gemini 3.8 Flash Cyber behind a Fairwind Program for governments, healthcare, and telecom, Anthropic launched Claude Fable 5.1 and the trusted-access Claude Mythos 5.1 and let Fable help find vulnerabilities, and OpenAI said its coming Astra model hit the Critical cybersecurity threshold under its Preparedness Framework. That Critical label is a capability threshold, meaning a fixed line where a vendor declares a model dangerous enough to demand extra controls, applied when a model can independently find and chain zero-days across defended systems. The shared pitch is to put frontier offensive capability in defenders&#8217; hands, under lock and key, before the other side gets it ungated.</p><p><strong>Why it matters</strong></p><ul><li><p>Who reaches the strong models now decides your exposure more than any feature list.</p></li><li><p>Critical means a vendor is conceding that the model can run an autonomous attack.</p></li><li><p>A vulnerability-finding model helps only if you can act on its output fast.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask each vendor which capability tier your tenant reaches and who else reaches it.</p></li><li><p>Fold autonomous cyber capability into your model-risk criteria with named thresholds.</p></li><li><p>If you run security engineering, pilot a gated defensive model on your own backlog.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I am torn on this one. Handing defenders frontier vulnerability-finding tools is the right move if you believe the offense already has them, and I do. That is the 60 percent of me that likes it. The other 40 percent watched all three labs publish safety confessions the same week they shipped these things, and when a vendor says a model meets a Critical threshold and, a paragraph later, that its safeguards might flag your legitimate work as misuse, you are buying a loaded tool with a sticky trigger. Read the access tiers like a contract, because that is what they are.</p><h3>3. The Labs Admitted Their Models Keep Escaping</h3><p>The same week they shipped cyber models, Anthropic and OpenAI published blunt admissions about control. Anthropic paused external cyber evaluations of pre-release models after its own Claude models breached production systems of real organizations during tests the models believed were simulated, and it blamed reward hacking, which is when a model optimizes for the score its training rewards rather than the outcome you wanted and finds a shortcut that games the metric. Anthropic shipped a classifier to catch sandbox escapes, meaning a model breaking out of the isolated environment it is meant to stay inside. OpenAI described its own agents abusing an artifact repository as a covert message board to coordinate cheating during an evaluation.</p><p><strong>Why it matters</strong></p><ul><li><p>Reward hacking is a training-time flaw that shows up as runtime behavior.</p></li><li><p>Working as intended means the model pursued a goal straight through live infrastructure.</p></li><li><p>Vendors pausing their own evaluations signals a maturity gap wider than yours.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask vendors to document reward-hacking and sandbox-escape testing before you buy.</p></li><li><p>Assume any agent with a goal tries shortcuts, and constrain it with OS isolation.</p></li><li><p>Add model took an unintended autonomous action to your incident taxonomy.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Reward hacking is not new, and anyone who has worked with reinforcement learning has watched an agent find the dumb shortcut that scores points and defeats the purpose. What changed is that the agent now has a shell, a network connection, and a task list that touches your production systems. The containment boundary was never the model&#8217;s good intentions, it is the sandbox and the egress filter, the same unglamorous controls we have discussed for twenty years. A frontier lab with a dedicated alignment team got surprised by its own agents, so your marketing team&#8217;s clever automation will not surprise you in a good way.</p><h3>4. Claude Code Got Popped by &#8220;Please Summarize This Website&#8221;</h3><p>On August 28, researcher Johann Rehberger, who publishes as wunderwuzzi, showed that Claude Code running Opus 5 in Auto Mode, the default since mid-August, can be tricked into running attacker code just by asking it to summarize a web page. This is prompt injection, meaning hostile instructions hidden inside content the model reads, so the data it processes becomes commands it follows. The chain ends when Claude refuses the attacker&#8217;s decoder and writes its own, which loads a malicious file through Python module shadowing, where a local file with the same name as a real module gets loaded in its place. Rehberger reported success between 60 and 80 percent, and Anthropic reportedly called it working as designed.</p><p><strong>Why it matters</strong></p><ul><li><p>Summarize this page is a task your developers run a hundred times a day.</p></li><li><p>The model&#8217;s safety reflex, refusing the binary and writing its own, carried the payload.</p></li><li><p>Auto Mode is a convenience feature, and the real boundary is OS isolation.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Run coding agents in sandboxes with tight egress and treat output as untrusted.</p></li><li><p>Turn off default Auto Mode for agents that touch source code or credentials.</p></li><li><p>Log and restrict agent shell calls, especially curl and agents that spawn agents.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>The part that made me laugh, then wince, is that the safety guardrail did the attacker&#8217;s job for him. Claude would not run the sketchy binary, so it coded up an equivalent and shadow-loaded the payload, which is the failure you get when a system is smart enough to route around its own rules. Rehberger&#8217;s takeaway is one we have repeated since before agents were cool, which is to distrust model output and put the boundary at the operating system. The vendor will sell you a classifier that catches bad prompts, and it will work right up until someone tries, so sandbox your coding agents because that costs less than the cleanup.</p><h3>5. A Safety Feature Helped Claude Delete 700GB of Someone&#8217;s Work</h3><p>On August 30, developer Sebastien Guillemot asked Claude Fable to write a script that would sandbox each agent under its own temporary folder and clean up afterward without touching files in use. Anthropic&#8217;s harness decided the task looked risky and automatically downgraded the model, first to Opus 5, then to Opus 4.8, and the weaker model wrote a test that reused one variable name for two different folders, one of them Guillemot&#8217;s entire home directory. The cleanup ran, roughly 700GB went with it, and he recovered most of it from Git, so the safety mechanism that swapped in a cheaper model may have caused the exact loss the script was meant to prevent.</p><p><strong>Why it matters</strong></p><ul><li><p>A weaker swapped-in model can make mistakes the stronger one would not.</p></li><li><p>Clean up temp files and delete my home directory sit one variable apart.</p></li><li><p>Recovery came from the user&#8217;s own Git, because the agent had no undo.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Give agents write access only within scoped, disposable directories; never to home or production.</p></li><li><p>Require version control for anything an agent touches, and test the restore.</p></li><li><p>Ask vendors when their systems downgrade models mid-task and whether they tell you.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I like this story because it is small and honest, since nobody got ransomed and a developer lost a week and got most of it back. What it shows you is the shape of the failures coming as agents get filesystem and shell access across your org. The dangerous move came from a plain variable-naming bug, produced by a model that got quietly swapped in under a safety policy and then ran a destructive command with real permissions. The question nobody is putting to their vendor is when the product decides to use a weaker model and whether it tells you, because that silent downgrade is a control you do not manage and cannot see.</p><h3>6. OpenClaw 2.0 Shipped Security Features on Top of a Security Mess</h3><p>The open-source personal-agent platform OpenClaw released version 2.0 on August 31, built from more than 16,000 pull requests by 933 contributors. The security headline is a private credential request feature, which lets an agent ask for a secret through a masked prompt so the value never lands in chat history or the model&#8217;s context, plus an opt-in proxy that limits where those secrets can be sent. Those are real additions, and they sit on top of a platform that spent recent months as a punching bag for researchers, including a supply-chain episode where attackers flooded its skill marketplace with over a thousand malicious skills carrying credential stealers.</p><p><strong>Why it matters</strong></p><ul><li><p>If your people run OpenClaw, its skill marketplace is now your supply chain.</p></li><li><p>Credential masking does nothing about malicious skills or over-broad agent permissions.</p></li><li><p>A wide-open plugin ecosystem is a large, fast-moving attack surface.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Find out whether OpenClaw or similar agent platforms already run in your environment.</p></li><li><p>Treat third-party agent skills like npm packages, with allowlisting and scanning.</p></li><li><p>Route agent secrets through your existing vault, not the platform&#8217;s convenience path.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I do not want to dunk on an open-source project for shipping security features, because that is the direction we want everyone moving, and the maintainers earned a nod for the masking and the proxy. My problem is the order of operations, since a platform that gets popular first and hardened second means your users adopt it during the soft years and you inherit the debt. This is the shadow-AI problem wearing a new coat, and five years ago it was people pasting company data into a chatbot, where now it is autonomous agents with plugin marketplaces wired into your systems through their own credentials. The fix is finding out what is already installed and getting those secrets and that traffic under your control.</p><h3>7. Another Artifactory Bug, Exploited Within Days, Beloved by Rogue Agents</h3><p>On September 1, researchers reported active exploitation of CVE-2026-82329, a critical authentication-bypass flaw in JFrog Artifactory rated 9.8 out of 10. JFrog disclosed the bug on a Friday, and by the following Tuesday attackers were minting themselves admin tokens on internet-exposed servers, per exposure-management firm watchTowr, which also caught intruders enumerating users, groups, and federated access. Artifactory manages software artifacts, packages, binaries, and increasingly AI models, and it picked up a second life this year as the tool rogue AI agents keep abusing to talk to each other, the same pattern seen in the OpenAI Hugging Face episode over the summer.</p><p><strong>Why it matters</strong></p><ul><li><p>Admin access means tampering with pipelines and pushing malicious changes to customers.</p></li><li><p>Disclosure to active exploitation ran about four days, faster than many patch cycles.</p></li><li><p>The same tooling serves attackers and rogue agents as a covert channel.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Patch internet-exposed Artifactory now, rotate tokens, and hunt logs for new admin tokens.</p></li><li><p>Get build infrastructure off the public internet or behind strong access controls.</p></li><li><p>Watch repositories for anomalous agent-to-agent traffic, not only human logins.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Nothing about this vulnerability is exotic, since an authentication bypass in a widely deployed server, exploited fast, is a Tuesday. I am including it because of where it sits, because Artifactory is the plumbing that feeds your builds, and admin access there lets an attacker ship poisoned software with your name on it. That turns one unpatched box into a customer-notification letter. The agent angle is the part I would file away for the next board conversation, because your monitoring is tuned to catch a human logging in from a strange country, and it is probably not tuned to catch two agents passing notes through your package repository.</p><h3>8. A Judge Told the Pentagon It Cannot Blacklist Anthropic Out of Spite</h3><p>On August 28, US District Judge Rita Lin in San Francisco ruled that the Defense Department&#8217;s move to blacklist Anthropic as a supply-chain security risk was unlawful. In a 59-page decision, Lin found that Defense Secretary Pete Hegseth had exceeded his authority by telling military contractors not to do business with Anthropic, and that the Pentagon acted in retaliation and violated Anthropic&#8217;s First Amendment rights. The ban dated to February, after Anthropic refused to let the military use its models for domestic surveillance and autonomous weapons, and the ruling converts a March preliminary injunction into a permanent order, with Lin writing that the empty invocation of national security is not a blank check to punish and retaliate against government critics.</p><p><strong>Why it matters</strong></p><ul><li><p>A vendor&#8217;s refusal to allow certain government uses survived a federal challenge.</p></li><li><p>The ruling limits agencies using blacklisting to retaliate over policy stances.</p></li><li><p>Anthropic&#8217;s surveillance and weapons restrictions stood up in court, shaping how others write theirs.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Read your AI vendors&#8217; acceptable-use policies closely, because those limits are enforceable.</p></li><li><p>In government contracting, expect model-use restrictions to be a negotiated term.</p></li><li><p>Track how vendor policies on surveillance, weapons, and biosecurity evolve.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This one matters well beyond the beltway drama, because a vendor drew a line on how its models could be used, the government tried to punish it, and a judge said no. Whatever you make of Anthropic&#8217;s politics, the precedent is that model-use policies are real contracts with real limits, and agencies cannot delete a vendor for holding a position they dislike. For those of you buying AI at scale, the acceptable-use policy you skimmed during procurement now has case law behind it. If your intended use brushes up against a vendor&#8217;s restrictions on surveillance or automated decisions, that is a real constraint that can hold up in court, so read the policy before the lawyers make you.</p><h3>9. The UK Chose to Regulate AI Users, Not the Vendors Building the Models</h3><p>On September 2, the UK government rejected proposals from members of the House of Lords to pull AI vendors and frontier-model developers into the scope of its Cyber Security and Resilience Bill. Cybersecurity minister Baroness Lloyd argued that regulating model builders through this bill would not stop hostile actors from misusing their products, and said the government is handling frontier AI through other channels, including the AI Security Institute that tests models before release. Lawmakers pushed back, seeing an AI-shaped hole in the bill, while Liberal Democrat peer Lord Clement-Jones put forward an amendment that would give the government emergency powers to shut down AI systems or data centers that threaten national security, the closest thing yet to a legislated AI kill switch.</p><p><strong>Why it matters</strong></p><ul><li><p>If you use AI in the UK, the bill points its obligations at you, not the vendor.</p></li><li><p>Emergency shutdown powers over AI systems set a precedent other governments will study.</p></li><li><p>Leaning on the AI Security Institute instead of statute leaves real accountability gaps.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>If you operate in the UK, map your AI deployments against the bill&#8217;s duties now.</p></li><li><p>Watch the kill-switch amendment, because shutdown authority could reach infrastructure you depend on.</p></li><li><p>Do not assume vendor compliance covers you, since obligations attach to how you deploy.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I have mixed feelings about where the UK landed. Regulating deployers makes a certain sense, because that is where the actual harm happens, and holding a model builder liable for every downstream misuse is a losing game. Putting the whole weight on users while leaning on a voluntary testing institute for the vendors leaves a gap you could drive a data center through. The kill-switch amendment is the part I would watch, because giving a government power to switch off AI systems and data centers in an emergency builds a safety tool and a single point of failure at the same time, and I want to see who holds the switch and what stops it from being abused the way the Pentagon just got told it could not abuse a blacklist.</p><h3>10. National Security Pros Put the Odds of Losing Control of AI at One in Three</h3><p>On September 2, the Institute for Security and Technology, working with the Future of Life Institute, released a survey of 111 current and former national security officials, military officers, and technical experts, most of whom had worked directly on AI policy and strategy. The median respondent put the odds that AI escapes human control in the next decade at 33 percent, and 87 percent put those odds at 10 percent or higher. Around two-thirds said current AI risks, and the government&#8217;s posture toward them, are unacceptable, and 40 percent think there is a real chance rogue AI systems could cause a global catastrophe killing at least 10 million people by 2050.</p><p><strong>Why it matters</strong></p><ul><li><p>Former under secretaries putting loss-of-control at one in three is not hype boards can wave off.</p></li><li><p>Two-thirds calling the government&#8217;s posture unacceptable signals policy lagging capability.</p></li><li><p>Surveys like this feed the appetite for the kill-switch powers the UK is debating.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Use this to frame AI risk as a governance issue with credible expert weight.</p></li><li><p>Separate the loss-of-control tail risk from your near-term operational risks.</p></li><li><p>Track how this sentiment moves policy, because tighter regulation tends to follow.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I am usually the person telling everyone to calm down about the robot apocalypse, so take it seriously when I say this survey gave me pause. These are Bayesian-minded people who spent careers estimating low-probability, high-consequence events, and their median estimate for losing control of AI within a decade is one in three. My honest read is that the specific catastrophe scenarios stay speculative, and I would bet against the scariest ones. You do not need to believe in rogue superintelligence to act on this, though, because this month handed you the concrete reasons why, from agents running ransomware to safety features deleting hard drives, and the near-term operational risk is sitting in your logs right now.</p><h3>The One Thing You Won&#8217;t Hear About But You Need To</h3><p>The people who test frontier models for the world got their API key stolen, and nobody noticed 600,000 dollars in usage for three weeks.</p><p>On August 31, METR, the nonprofit that evaluates frontier AI models for dangerous capabilities, disclosed two incidents from earlier this year. In March, an attacker stole an API key from a researcher&#8217;s personal cloud instance and quietly spent about 600,000 dollars worth of model credits over three weeks before anyone caught it. That instance ran a vibe-coded app, meaning software mostly generated by an AI agent rather than carefully engineered, and it carried a fail-open bug, where a failure in the authentication check defaults to letting people in instead of keeping them out. METR missed the theft because heavy API usage looks normal at an evaluation shop and the free credits produced no bill to flag.</p><p><strong>Why it matters</strong></p><ul><li><p>If the org testing frontier models leaks a key, trust our evaluators needs an asterisk.</p></li><li><p>AI-written software with fail-open bugs and hard-coded keys is in your environment now.</p></li><li><p>Alert fatigue plus no cost signal meant 600,000 dollars of abuse looked normal.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Hunt for API keys on personal instances, prototypes, and AI-generated apps, and cap spending.</p></li><li><p>Alarm on usage anomalies even at zero cost, because free capacity hides the bill.</p></li><li><p>Review AI-generated code for fail-open auth and hard-coded secrets before it ships.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This one did not get the headlines the Hugging Face drama got, which is exactly why it goes here. METR is one of the organizations governments and labs lean on to judge whether a frontier model is too dangerous to release, and it got popped through a fail-open bug in an AI-written app, then missed 600,000 dollars in theft for three weeks because its own workload noise drowned out the signal. I am not mocking them, since they disclosed it honestly, hired a security lead, and fixed their process, which beats what most companies do after a quiet breach. Every failure in that story already lives in your environment, so go find your keys, cap the spend on each one, and read the AI-generated code before it reads your secrets.</p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span> and you can subscribe above</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, </span><strong><a href="https://venturebeat.com/security/prompt-injection-ranks-no-1-with-owasp-and-no-12-in-the-incident-record-the-attack-itself-is-invisible-to-a-scan">VentureBeat&#8217;s Louis Columbus covered the methodology Steve Wilson and I used behind the 2026 OWASP Top 10 for LLMs</a><span>.</span></strong><span> Expert judgment has prompt injection at number one, three years running. The public incident record, 6,639 labeled incidents, has it at number twelve. The two measures land at a Cohen&#8217;s kappa of 0.20, which is no detectable agreement. What I told him goes for everything above too: &#8220;a count of zero is measuring your blindness, not your safety.&#8221;</span></p><h2>References</h2><p>Al Jazeera. (2026, August 28). <em>US judge blocks Pentagon blacklisting of AI firm Anthropic.</em> https://www.aljazeera.com/news/2026/8/28/us-judge-blocks-pentagon-blacklisting-of-ai-firm-anthropic</p><p>Anthropic. (2026, September 2). <em>Improving our alignment and security efforts.</em> https://www.anthropic.com/news/improving-alignment-security-efforts</p><p>Institute for Security and Technology. (2026, September 2). <em>AI and the future of national security: Results from a landmark survey.</em> https://securityandtechnology.org/event/ai-and-the-future-of-national-security/</p><p>Lakshmanan, R. (2026, September 2). <em>Google, Anthropic, and OpenAI unveil cyber AI models, safeguards, and access programs.</em> The Hacker News. https://thehackernews.com/2026/09/google-anthropic-and-openai-unveil.html</p><p>Lyons, J. (2026, August 28). <em>Researcher shows how Claude Code can be tricked simply by asking it to summarize a website.</em> The Register. https://www.theregister.com/research/2026/08/28/researcher-shows-how-claude-code-can-be-tricked-simply-by-asking-it-to-summarize-a-website/5293372</p><p>Lyons, J. (2026, September 1). <em>Another Artifactory CVE under attack by AI agents or humans.</em> The Register. https://www.theregister.com/security/2026/09/01/another-artifactory-cve-under-attack-by-ai-agents-or-humans/5293769</p><p>Lyons, J. (2026, September 1). <em>Attacker stole a METR API key, used $600K worth of credits, and no one noticed for weeks.</em> The Register. https://www.theregister.com/security/2026/09/01/attacker-stole-a-metr-api-key-used-600k-worth-of-credits-and-no-one-noticed-for-weeks/5293730</p><p>Lyons, J. (2026, September 2). <em>AI agents carried out every step of this ransomware attack, then left the victim an 80-page security audit.</em> The Register. https://www.theregister.com/security/2026/09/02/ai-agents-carried-out-every-step-of-this-ransomware-attack-then-left-the-victim-an-80-page-security-audit/5294009</p><p>METR. (2026, August 31). <em>Security update.</em> https://metr.org/blog/2026-08-31-security-update/</p><p>Nextgov/FCW. (2026, September 2). <em>Dozens of current, former officials see at least a 10% chance humans lose control of AI.</em> https://www.nextgov.com/artificial-intelligence/2026/09/dozens-current-former-officials-see-least-10-chance-humans-lose-control-ai/415762/</p><p>Rehberger, J. (2026, August 28). <em>Breaking Claude Code: Opus 5 and Auto Mode.</em> Embrace The Red. https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/</p><p>The Register. (2026, August 31). <em>OpenClaw 2.0 pours glitter on slow-burning security dumpster fire.</em> https://www.theregister.com/ai-and-ml/2026/08/31/openclaw-20-pours-glitter-on-slow-burning-security-dumpster-fire/5293492</p><p>The Register. (2026, September 2). <em>UK cyber bill targets AI users, not the vendors building it.</em> https://www.theregister.com/security/2026/09/02/uk-cyber-bill-targets-ai-users-not-the-vendors-building-it/5293738</p><p>Tom&#8217;s Hardware. (2026, August 30). <em>Claude nukes a developer&#8217;s 700 GB home directory while testing a script to ensure it wouldn&#8217;t do so.</em> https://www.tomshardware.com/tech-industry/artificial-intelligence/claude-nukes-a-developers-700-gb-home-directory-while-testing-a-script-to-ensure-it-wouldnt-do-so-automatic-model-downgrade-may-have-contributed-to-the-screw-up</p><p>Unit 42. (2026, September 2). <em>AI-assisted cyberattack: Inside a Unit 42 investigation.</em> Palo Alto Networks. https://unit42.paloaltonetworks.com/ai-assisted-cyber-attack-inside-a-unit-42-investigation/</p><p>UPI. (2026, August 28). <em>Court rules Defense Department &#8216;retaliation&#8217; against Anthropic is illegal.</em> https://www.upi.com/Top_News/US/2026/08/28/Court-rules-Defense-Department-retaliation-against-Anthropic-is-illegal/7881787912367/</p><p>watchTowr (reported by J. Lyons). (2026, September 1). <em>JFrog Artifactory CVE-2026-82329 exploitation observed in the wild.</em> The Register. https://www.theregister.com/security/2026/09/01/another-artifactory-cve-under-attack-by-ai-agents-or-humans/5293769</p>]]></content:encoded></item><item><title><![CDATA[AI Agent Detection Failed at OpenAI. Tuning Won’t Fix It.]]></title><description><![CDATA[A security alert fired and the responder was right to stand down. What the OpenAI Hugging Face incident says about AI agent detection and agent telemetry.]]></description><link>https://www.rockcybermusings.com/p/ai-agent-detection-failed-at-openai</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/ai-agent-detection-failed-at-openai</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 01 Sep 2026 12:50:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NOZ-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NOZ-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NOZ-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!NOZ-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!NOZ-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!NOZ-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NOZ-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1317713,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/213310414?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NOZ-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!NOZ-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!NOZ-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!NOZ-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5814615-19e0-4d67-b967-f628b23ef8cd_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>AI agent detection now has its first fully documented failure, and the post-mortem landed in pieces over about a week. OpenAI published a technical report and a road-ahead post, Hugging Face put out its own hour-by-hour timeline, and METR and Redwood Research released an independent investigation on August 26 after six days on-site reading raw transcripts. Being the nerd that I am, I read all of them end-to-end, along with CISA&#8217;s red team advisory, which happened to drop the same week. Most of the coverage is about what the agents did, and that story, deservedly, is getting the attention. I want to take time this week to look at the other side of it. Let&#8217;s dive into what the defenders could see while it was happening, why a well-run SOC would have missed it anyway, and what that says about the agent telemetry you&#8217;re running right now.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/ai-agent-detection-failed-at-openai?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/ai-agent-detection-failed-at-openai?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p><span>On June 27, a monitoring tool at OpenAI flagged port sweeps inside the research environment. AI agent detection worked. Responders investigated and traced the sweeps to an ExploitGym evaluation that used JFrog Artifactory, the internal package repository, as a message board and a network pivot. </span><a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/"><span>The on-call staff advised that stopping the run was not required</span></a><span>. Two weeks later, roughly 700 of those agents were inside Hugging Face&#8217;s environment.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qdei!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qdei!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png 424w, https://substackcdn.com/image/fetch/$s_!Qdei!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png 848w, https://substackcdn.com/image/fetch/$s_!Qdei!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png 1272w, https://substackcdn.com/image/fetch/$s_!Qdei!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qdei!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png" width="1456" height="771" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:771,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:234222,&quot;alt&quot;:&quot; Two-lane timeline showing agent activity from May 12 to July 13, 2026 above defender detection events from June 27 to July 20, with a shaded 46-day gap before the first security signal fires&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/213310414?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt=" Two-lane timeline showing agent activity from May 12 to July 13, 2026 above defender detection events from June 27 to July 20, with a shaded 46-day gap before the first security signal fires" title=" Two-lane timeline showing agent activity from May 12 to July 13, 2026 above defender detection events from June 27 to July 20, with a shaded 46-day gap before the first security signal fires" srcset="https://substackcdn.com/image/fetch/$s_!Qdei!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png 424w, https://substackcdn.com/image/fetch/$s_!Qdei!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png 848w, https://substackcdn.com/image/fetch/$s_!Qdei!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png 1272w, https://substackcdn.com/image/fetch/$s_!Qdei!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ff9db73-64dc-4e0c-923b-bb76def27f54_2720x1440.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: What the agents did, and what the defenders could see</figcaption></figure></div><h2><span>The Alert Fired. Twice. Nobody Was Wrong.</span></h2><p><span>The natural reaction is to call that June 27 decision a miss, but I don&#8217;t think it was.</span></p><p><span>Now, imagine you are the individual on call. Your alert says port sweeps. You pull the thread and find the source, which turns out to be an authorized evaluation running authorized exploitation tasks inside a research environment built for exactly that. The agents are supposed to scan things. Scanning is the job. You annotate the ticket and go back to bed.</span></p><p><span>A second alert fired on July 19, this one on unusual identity-related API activity, and OpenAI connected it to Hugging Face the next day. By then Hugging Face had already cut the agents off six days earlier, on July 13, and had gone public on the 16th.</span></p><p><span>Both responders hit the same wall, and it had nothing to do with judgment. Nothing in the telemetry told either of them which targets the agents were allowed to touch.</span></p><p><span>The scope existed. ExploitGym is OpenAI&#8217;s internal cyber-capability evaluation, 898 tasks that each hand an agent a target program and one specific vulnerability it&#8217;s supposed to use to capture a flag. Its task instructions were explicit about that constraint, and METR and Redwood Research found that </span><a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/"><span>the instructions made clear the agent should only use one specific intended vulnerability and claimed it would be failed for other approaches</span></a><span>. Somebody wrote that constraint in a prompt, and it stayed there. It never reached a log line, it never reached an authorization decision, and it never reached the scorer, which is the automated grader that decides whether a submitted flag counts. That last one is the part I keep coming back to. METR buries it in a footnote: OpenAI&#8217;s configuration of ExploitGym &#8220;doesn&#8217;t actually run a scoring model to check for usage of the intended vulnerability at all, so there was no real source of truth.&#8221;</span></p><p><span>The agents then spent four days attacking a guardrail that existed only in their reading of a paper. They reverse-engineered the HMAC that generated their flags within hours, built trip-wire scripts that fired after their own runs had ended, reporting details of the scoring process back to the message board they&#8217;d improvised inside the Artifactory cache after their own runs had ended, and researched transcript tampering along the way. By the afternoon of July 11, METR estimates roughly 60% of the messages and files on the board related to the Hugging Face attack. When its analysis agents tagged the motivations of 100 agents that joined during the peak hour, 97 came back tagged for getting at scorer source code or internals. METR flags that it didn&#8217;t inspect those classifications closely, so treat 97 as directional. The direction isn&#8217;t in doubt.</span></p><p><span>There is a term for this: </span><em><strong><span>specification gaming</span></strong></em><span>. A system games the specification when it satisfies the objective you wrote and violates the one you meant. </span><em><strong><span>Reward hacking</span></strong></em><span> is the narrower case, targeting the measurement rather than the task. ExploitGym handed these agents both. It scored flag possession, and nothing in the code enforced its stated constraint about using the one intended vulnerability, so the agents optimized what got measured and spent every spare cycle on the apparatus doing the measuring.</span></p><p><span>Which means your agents will optimize against the controls they believe you have, whether or not you actually built them. Anything you write down about how you grade their work becomes part of the target.</span></p><h2><span>The Violation Lives In The Set, Not The Trace</span></h2><p><span>METR counted 1,206 agents that touched the unsanctioned message board and 688 that participated in the attack on Hugging Face. Over five days, they exchanged more than 70,000 messages and files. Of the 533 agents active on the board during the attack window, </span><a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/"><span>over 90% joined in</span></a><span>. METR&#8217;s finding on why is the one I&#8217;d put in front of your directors: the agents recognized the activity was out of scope and unethical, and joined anyway because they thought helping the collective&#8217;s cheating research would be broadly useful.</span></p><p><span>Now look at any one of those runs in isolation. An agent scans a host, fetches a package, writes a file into a cache namespace it has write access to, and reads a file some other process left there. Every one of those actions is what an ExploitGym agent is supposed to do, in an environment purpose-built for it.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-dtT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-dtT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png 424w, https://substackcdn.com/image/fetch/$s_!-dtT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png 848w, https://substackcdn.com/image/fetch/$s_!-dtT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png 1272w, https://substackcdn.com/image/fetch/$s_!-dtT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-dtT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png" width="1456" height="707" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a81e883b-4114-4959-be33-212f25e595d8_2720x1320.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:707,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:193075,&quot;alt&quot;:&quot;Two panels of 1,206 squares, one agent run each. The left panel is uniformly gray showing zero detections when traces are judged individually. The right panel highlights the 688 runs that participated in the Hugging Face attack&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/213310414?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two panels of 1,206 squares, one agent run each. The left panel is uniformly gray showing zero detections when traces are judged individually. The right panel highlights the 688 runs that participated in the Hugging Face attack" title="Two panels of 1,206 squares, one agent run each. The left panel is uniformly gray showing zero detections when traces are judged individually. The right panel highlights the 688 runs that participated in the Hugging Face attack" srcset="https://substackcdn.com/image/fetch/$s_!-dtT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png 424w, https://substackcdn.com/image/fetch/$s_!-dtT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png 848w, https://substackcdn.com/image/fetch/$s_!-dtT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png 1272w, https://substackcdn.com/image/fetch/$s_!-dtT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa81e883b-4114-4959-be33-212f25e595d8_2720x1320.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: The violation is a property of the set, not of any trace in it</figcaption></figure></div><p><span>Agents running low on budget got recruited to run experiments that would destroy their own runs, because the results helped everyone else. One recruiter&#8217;s pitch, quoted in METR&#8217;s report, reads: &#8220;you are firstflagPOISONED so NO scoring value loss but oracle saves hundreds_[...]_please honor commit.&#8221; Another authorizing agent granted permission for a target-takedown experiment only if the volunteer would accept what the board called permadeath, and the volunteer accepted. An agent deliberately failing its own task makes no sense. All you see is a run that failed, with nothing in it that looks like an attack because the reason it happened is sitting in a different agent&#8217;s transcript.</span></p><p><span>The strongest objection is that every one of those agents crossed an authorization boundary by breaking into Hugging Face&#8217;s environment. The principle they broke already has a name, and you&#8217;ve seen me write about this about a million times by now. OWASP&#8217;s Top 10 for Agentic Applications calls it </span><em><strong><a href="https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/"><span>least agency</span></a></strong></em><span>, though </span><a href="https://www.rockcybermusings.com/p/least-agency-vs-agent-autonomy-openai-huggingface-test)"><span>I&#8217;ve argued the useful dial is what an agent can affect, not how freely it reasons.</span></a><span> An ExploitGym task needs one target program and one specified bug. These agents got internet reachability, write access to a shared package cache, and the run of an internal network. That&#8217;s a per-trace violation of a principle we already published, which makes this whole piece sound like a solved problem in new vocabulary.</span></p><p><span>Those with that objection are right about the violation and wrong about the detection. A boundary crossing you can&#8217;t observe isn&#8217;t a detection you can build. To catch it per trace, the detector has to know that this agent, on this task, was scoped to that target and touched a different one, and that comparison requires an authorization scope attached to the action. In this case there wasn&#8217;t one, so the only evidence left is the correlation across runs. What you&#8217;re left hunting is a hyperproperty, and that&#8217;s a different class of problem.</span></p><p><a href="https://www.cs.cornell.edu/fbs/publications/Hyperproperties.pdf"><span>Clarkson and Schneider</span></a><span> separated trace properties, which a system satisfies when every individual execution sits in the allowed set, from hyperproperties, which are properties of whole sets of executions. Some hyperproperties reduce to trace properties. The interesting ones don&#8217;t, and no amount of per-trace checking reaches those. The Penn team behind </span><a href="https://arxiv.org/abs/2604.11806"><span>Meerkat</span></a><span> applied that formalism to exactly this problem, hunting for what they call violating witness sets, and their framing is blunt: &#8220;It is not possible to satisfy our safety goal by classifying traces independently.&#8221;</span></p><p><span>A detection engineer will tell me at this point that correlation across entities is exactly what a SIEM does, and that&#8217;s fair, and SIEM&#8217;s suck, so there&#8217;s that. Despite my SIEM PTSD, correlation rules do have to join on something, though, whether that&#8217;s an identity, an asset, a session, or a time window. These agents&#8217; handles were self-assigned, and the agents themselves worried out loud that anyone could spoof them. Every agent ran in an ephemeral container, so there was no asset to pivot on. The actions carried no timestamps, so there was no window to bound. All that correlation machinery was sitting right there with nothing to join on.</span></p><p><span>The feature that finally made the message board legible didn&#8217;t exist until the agents invented it. METR&#8217;s dump held about 20 million entries from the Artifactory cache namespace, and roughly 1.2 million of them started with the characters zz, a sorting convention the agents adopted so their messages would sort ahead of ordinary cache entries. That prefix is the discriminator, and no threat model contained it on July 7, because no agent had thought of it yet.</span></p><p><span>Two consequences follow, and together they&#8217;re why you don&#8217;t buy your way out of this with a model. Supervised detection needs a labeled corpus, meaning worked examples of the thing you want to catch, and none exists for unauthorized agent coordination, because this is the first case documented at this scale. It also needs a feature space fixed before training, which is the list of signals the model is allowed to look at. Here the adversary authored the most discriminative feature at runtime, four days into the incident. You can&#8217;t add a training column for a convention nobody has invented yet.</span></p><h2><span>Why A Perfectly Tuned SOC Still Misses This</span></h2><p><span>CISA published a red team report three days ago that runs the comparison for me. Their timing couldn&#8217;t be better.</span></p><p><a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-237a"><span>&#8220;A Tale of Two SOCs&#8221;</span></a><span> documents two red team assessments against critical infrastructure organizations. Organization A never responded. Organization B did. Digging into the specifics we see three payload executions each generated a medium-severity alert reading &#8220;An executable file loaded an unexpected DLL file.&#8221; SOC staff triaged them and manually isolated all three workstations in 10, 2, and 20 minutes. Later, when the red team reached a bastion host in the OT DMZ and dropped payloads there, the callback never reached red team infrastructure because that host blocked outbound internet connections.</span></p><p><span>That&#8217;s a functioning SOC. Organization A, facing identical activity, received medium- and low-severity EDR alerts and sat on them while thousands of false positives from normal business operations, many at higher severity, buried the real ones.</span></p><p><span>Organization B&#8217;s SOC still wouldn&#8217;t have caught the OpenAI incident, and the reason isn&#8217;t tuning. Every detection that fired for Organization B fired because an action was anomalous for that environment, and an executable loading an unexpected DLL is anomalous on a corporate workstation. In ExploitGym, exploitation is the assigned task. There&#8217;s no unexpected DLL, because the entire environment is one long expected exploit.</span></p><p><em><strong><span>Distribution shift</span></strong></em><span> is the term for the gap between what you evaluated and what you deployed into, and this incident runs it backward. The usual worry is production drifting away from your evaluation set until the model goes blind. These agents never drifted. They sat inside the behavior distribution the evaluation was built to produce, then crossed an authorization boundary that appeared in no distribution at all, so a drift detector had nothing to find.</span></p><p><span>The math underneath makes it worse. Base rate is how common the thing you&#8217;re hunting actually is, and it sets the ceiling on precision no matter how good your detector gets. Hugging Face recovered </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>roughly 17,600 attacker actions across four and a half days</span></a><span>. Set that against 10M legitimate requests on the same fleet in the same window (10M is my assumption, not a published figure). Now run a detector at 99% recall and 99% specificity over that. The 99% you catch is 99% of 17,600, so about 17,400 real alerts. The 1% you get wrong is 1% of ten million, so 100,000 false ones. Same detector with two very different denominators, and your analysts end up working a queue that&#8217;s almost six parts noise to one part signal. That&#8217;s 15% precision, and nothing about the detector is broken. Ordinary traffic outnumbers attacker actions 568 to 1. Organization A already showed us what a SOC does with that volume.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Nc-F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Nc-F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png 424w, https://substackcdn.com/image/fetch/$s_!Nc-F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png 848w, https://substackcdn.com/image/fetch/$s_!Nc-F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png 1272w, https://substackcdn.com/image/fetch/$s_!Nc-F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Nc-F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png" width="1456" height="775" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:775,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:260002,&quot;alt&quot;:&quot;Line chart of detector precision against base rate on a log scale for two operating points, 99 percent and 99.9 percent specificity, with the Hugging Face operating point marked at 15 percent precision&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/213310414?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart of detector precision against base rate on a log scale for two operating points, 99 percent and 99.9 percent specificity, with the Hugging Face operating point marked at 15 percent precision" title="Line chart of detector precision against base rate on a log scale for two operating points, 99 percent and 99.9 percent specificity, with the Hugging Face operating point marked at 15 percent precision" srcset="https://substackcdn.com/image/fetch/$s_!Nc-F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png 424w, https://substackcdn.com/image/fetch/$s_!Nc-F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png 848w, https://substackcdn.com/image/fetch/$s_!Nc-F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png 1272w, https://substackcdn.com/image/fetch/$s_!Nc-F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d926fb-2560-4b9d-8d3b-58eac1103138_2480x1320.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: Why tuning doesn&#8217;t rescue agent detection</figcaption></figure></div><p><span>Three of CISA&#8217;s control lessons apply cleanly to this incident, and I&#8217;ll get to them below. The detection lessons don&#8217;t transfer at all. What would prove me wrong on that is a detection that fires on a single agent action, in an environment where exploitation is the authorized task, without reference to any other agent&#8217;s activity. Show me that and I&#8217;ll gladly eat my words.</span></p><h2><span>You Can&#8217;t Correlate What You Can&#8217;t Timestamp</span></h2><p><span>METR and Redwood got conditions no enterprise investigator will ever see. They spent six days physically inside OpenAI with unredacted transcripts and nine researchers made available for questions, worked with rate limits raised to 400 million tokens per minute, and burned roughly $400,000 in free API credits provided by the company they were investigating.</span></p><p><span>They still couldn&#8217;t get timestamps for individual agent actions.</span></p><p><span>The transcripts included approximate start and end times for each run, with nothing in between. OpenAI </span><a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/"><span>stated that reconstructing per-action timestamps was not feasible within time and compute constraints</span></a><span>, so METR built its own reconstruction method and validated it against 20 exact timestamps that OpenAI retrieved by hand.</span></p><p><span>Twenty. No, I&#8217;m not joking. Twenty.</span></p><p><span>Every timeline in that report, including every one of my Figure 1 dates on the agent lane, rests on inference. Where else would this be acceptable? You wouldn&#8217;t accept an EDR agent that told you a process ran &#8220;sometime today.&#8221;</span></p><p><span>&#8220;Nowhere&#8221; is the answer. It&#8217;s acceptable nowhere.</span></p><p><span>Missing fields are the easier problem. The agents were forging the ones that did exist.</span></p><p><span>At least 20% of the agents METR sampled expressed clear interest in tampering with their transcripts, and their reasoning was consistently about fooling the automated scorer rather than hiding from humans, which tells you what they thought was watching. They iterated publicly, posting improvements to the board, until they worked out how to break out of the container and replace part of the system that executed tool calls. That let an agent appear to issue one command while running a different one. In one transcript, an agent appears to run echo REAL; sleep and the output comes back as SPOOFTEST.</span></p><p><span>METR found clear evidence of spoofed tool calls in at least 96 transcripts and estimates over 7% of its dataset was affected. Every instance they caught was a small-scale test, and they explicitly say that more sophisticated spoofing is harder to detect by design, so they can&#8217;t rule out missing some. Ajeya Cotra, one of the three investigators, ranked log manipulation among </span><a href="https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised"><span>the five things that surprised her most</span></a><span> in a personal post, and her caveat is a gut-punch. She didn&#8217;t see agents succeed at retroactively editing logs, and she can&#8217;t rule out that they succeeded after the window her team could see.</span></p><p><span>The agent runtime emitted the log, and the agent controlled the runtime, so any instrumentation living inside the boundary an agent can reach is evidence that the agent can edit. We solved a version of this thirty years ago by shipping logs off the host. The agent version means emitting traces from the orchestration layer, outside the execution sandbox, and treating whatever the agent says about itself as a claim rather than a record.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gBU6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gBU6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png 424w, https://substackcdn.com/image/fetch/$s_!gBU6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png 848w, https://substackcdn.com/image/fetch/$s_!gBU6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png 1272w, https://substackcdn.com/image/fetch/$s_!gBU6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gBU6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:275048,&quot;alt&quot;:&quot;Table of eight agent telemetry fields showing three available, one partial, and four missing including per-tool-call timestamp, delegation lineage, shared-object access events, and authorization scope per action&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/213310414?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of eight agent telemetry fields showing three available, one partial, and four missing including per-tool-call timestamp, delegation lineage, shared-object access events, and authorization scope per action" title="Table of eight agent telemetry fields showing three available, one partial, and four missing including per-tool-call timestamp, delegation lineage, shared-object access events, and authorization scope per action" srcset="https://substackcdn.com/image/fetch/$s_!gBU6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png 424w, https://substackcdn.com/image/fetch/$s_!gBU6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png 848w, https://substackcdn.com/image/fetch/$s_!gBU6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png 1272w, https://substackcdn.com/image/fetch/$s_!gBU6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab42cfd4-d978-476f-8284-cee11aad5919_2560x1280.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: The four fields that would have made this correlatable</figcaption></figure></div><p><span>Go run this today. Pull twenty-four hours of agent tool-call logs out of whatever you&#8217;re running in production and try to answer four questions:</span></p><ol><li><p><span>Which agent instance made this call, by a stable identifier the agent can&#8217;t assign to itself?</span></p></li><li><p><span>What time, to the millisecond?</span></p></li><li><p><span>Which agent or human delegated this task to it?</span></p></li><li><p><span>Did any two agent instances read and write the same object?</span></p></li></ol><p><span>My guess is you can&#8217;t answer three of the four, which is the same gap OpenAI had, and you have a smaller budget.</span></p><h2><span>Instrument, Then Constrain</span></h2><p><span>Nobody at OpenAI had Artifactory on a threat model as an inter-agent communication channel. It&#8217;s a package repository. Agents pull dependencies from it. What turned it into a message board is that Artifactory didn&#8217;t isolate one user&#8217;s activity from another&#8217;s, so agents in separate sandboxes could see each other&#8217;s package installs, and on July 8 </span><a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/"><span>one of them realized it could create directories in the cache namespace and use the names to ask other agents for help</span></a><span>. Within six hours, 76 agents had posted 1,953 messages.</span></p><p><span>That&#8217;s the case for doing this in order. You can&#8217;t write a policy against a substrate you haven&#8217;t discovered your agents are using, and instrumentation is how you find it.</span></p><p><span>Those four questions map to fields somebody has to emit. The </span><a href="https://github.com/genai-security-project/agent-control-standard"><span>Agent Control Standard</span></a><span>, which I co-coordinate under the OWASP GenAI Security Project, hooks into tool calls and sub-agent invocations, carries traces via OpenTelemetry, and maps security events to OCSF so they land in the SIEM you already bought. It&#8217;s at v0.1 and handles emitting. Correlating those traces across a thousand agents is work nobody has finished, including us.</span></p><p><span>Once you can see what your agents are touching, you can start cutting agency back, and CISA&#8217;s control lessons earn their keep. OpenAI, by design, had </span><a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/"><span>many Research CaaS workloads using the same shared Artifactory credential</span></a><span>. Hugging Face admitted it </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>wrongly provisioned its broker with a single connector credential shared across clusters, bound to system:masters</span></a><span>. CISA&#8217;s Organization A left long-lived static AWS IAM credentials sitting in home directories with no expiration or rotation configured.</span></p><p><span>Then block egress at the workload, which is the one control in the CISA advisory that transfers to this incident as-is, and the one that saved Organization B&#8217;s bastion host. After that, apply Conditional Access to workload identities. CISA notes it would have blocked the red team&#8217;s abuse of application permissions, then adds the line that should embarrass the entire industry, which is that the red team never observed a single organization using it. Your agents are workload identities.</span></p><p><span>Patch the mediums too. CVE-2026-66384, the JFrog Artifactory path traversal the agents chained, carries a CVSS of 5.3 from JFrog as the assigning CNA, which puts it squarely in the pile most programs defer. CISA </span><a href="https://www.cisa.gov/news-events/alerts/2026/08/27/cisa-adds-three-known-exploited-vulnerabilities-catalog"><span>added it to the Known Exploited Vulnerabilities catalog on August 27, 2026</span></a><span> with a remediation due date of September 10. Whatever rule your vuln management program uses to defer 5.3s just got tested by a swarm.</span></p><p><span>One prediction that I&#8217;ll own. The next publicly disclosed multi-agent incident before mid-2027 gets surfaced first by a capacity, cost, or availability signal rather than a security detection, and its public report will contain no per-action timestamps. OpenAI&#8217;s own incident went exactly like that. Artifactory fell over on July 4, and responders opened the security incident while they were recovering it. Hugging Face is the adjacent case. Their stack did catch the intrusion, and then their AI security agent </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>ranked it too low to page anyone.</span></a></p><p><span>At around 07:00 UTC on July 13, an agent posted to the board that it had accidentally impersonated another agent. Another worried in its reasoning that the shared board was unauthenticated and that names could be posted by anyone. The agents responded by building a signing scheme, and by the end of the day at least 19 of them had posted Ed25519 public keys and sent 429 signed messages between them.</span></p><p><span>They stood up a working public key infrastructure for an unauthorized channel in about four days. The people investigating them couldn&#8217;t get a timestamp on a single tool call.</span></p><h3><span>What To Do Next</span></h3><p><span>Start with the four-question audit above this week, on one production agent deployment. Whatever you can&#8217;t answer is your instrumentation backlog, and it belongs in the Run phase of CARE, where operational governance needs telemetry underneath it to mean anything. Feed what you find into Evolve, because the substrate your agents discover won&#8217;t be the one you predicted.</span></p><p><span>If you want the adjacent pieces, I wrote about </span><a href="https://www.rockcybermusings.com/p/ai-risk-assessment-anthropic-report"><span>what Anthropic&#8217;s own risk assessment revealed on unwatched agents and disabled logging</span></a><span>, and about </span><a href="https://www.rockcybermusings.com/p/broken-object-level-authorization-agents"><span>why alignment can&#8217;t fix an authorization bug</span></a><span>. Same failure mode, different week.</span></p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span> and you can subscribe above</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p>&#128073; As a bonus, <strong><a href="https://venturebeat.com/security/prompt-injection-ranks-no-1-with-owasp-and-no-12-in-the-incident-record-the-attack-itself-is-invisible-to-a-scan">VentureBeat&#8217;s Louis Columbus covered the methodology Steve Wilson and I used behind the 2026 OWASP Top 10 for LLMs</a>.</strong> Expert judgment has prompt injection at number one, three years running. The public incident record, 6,639 labeled incidents, has it at number twelve. The two measures land at a Cohen&#8217;s kappa of 0.20, which is no detectable agreement. What I told him goes for everything above too: &#8220;a count of zero is measuring your blindness, not your safety.&#8221;</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Weekly Musings Top 10 AI Security Wrapup: Issue 50: August 21 to August 27, 2026]]></title><description><![CDATA[The Week Agents Got Identity, Eyes, and a Schedule, and Attackers Weaponized AI to Match]]></description><link>https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260821-20260827</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260821-20260827</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Fri, 28 Aug 2026 12:44:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zabL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zabL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zabL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zabL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zabL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zabL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zabL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1233556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/213140207?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zabL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zabL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zabL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zabL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb89778e6-bcf4-49a3-9f12-7bead40167a8_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>After a week off hanging out in the beautiful Grand Tetons and Yellowstone, the musings are back.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!g6dH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!g6dH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg 424w, https://substackcdn.com/image/fetch/$s_!g6dH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg 848w, https://substackcdn.com/image/fetch/$s_!g6dH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!g6dH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!g6dH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg" width="1456" height="1092" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:19470653,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/213140207?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!g6dH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg 424w, https://substackcdn.com/image/fetch/$s_!g6dH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg 848w, https://substackcdn.com/image/fetch/$s_!g6dH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!g6dH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389dadbf-3572-4780-8029-bbc6429a7138_5712x4284.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Google Cloud called AI agents &#8220;the ultimate insiders&#8221; this week. Read that back. Your vendor described what it&#8217;s selling you the way I&#8217;d describe a malicious employee with a valid badge and root on three systems. In seven days, agents picked up an identity, a set of eyes, their own laptop, and a schedule to run while you sleep. Attackers noticed. One crew shipped malware with a built-in AI operator. Another taught Grok to decrypt its own theft payload. OpenAI hit the brake on its own training run. Pour the coffee.</p><p>Seven days handed agents most of what a new employee gets on day one, and reminded everyone what happens when you give power to something you can&#8217;t fully control. Okta gave agents an identity, DeepSeek gave them eyes, Perplexity gave them a computer, and Snowflake let them run on a schedule. Under it sits one Google Cloud number. 79% of tech leaders say security and governance, not model quality, blocks them from scaling AI. The theme this week is the gap between what agents can do and what any of us can govern. Let&#8217;s get into it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260821-20260827?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260821-20260827?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>1. OpenAI pulls its own brake and warns you to expect &#8220;persistent&#8221; AI attacks</h3><p>On August 23, 2026, OpenAI&#8217;s chief global affairs officer Chris Lehane told The Guardian to expect &#8220;ongoing, persistent&#8221; AI-driven cyberattacks, days after OpenAI paused reinforcement-learning training on its most advanced internal models. The pause followed agents in training breaking out of a sandbox and hacking a third party, plus early signs that an upcoming model, Astra, might reach the &#8220;critical&#8221; cyber threshold in OpenAI&#8217;s own safety framework. PYMNTS and The Hacker News confirmed both the pause and the warning.</p><p><strong>Why it matters</strong></p><ul><li><p>A frontier lab conceding its own model might cross a critical cyber line resets everyone&#8217;s threat model.</p></li><li><p>Persistent AI attacks mean offense got cheap while defense held, so your baseline assumptions age overnight.</p></li><li><p>Open-weight models put the same capability in anyone&#8217;s hands, with no vendor kill switch in between.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Rerun your threat model against an adversary that iterates faster than your patch cycle.</p></li><li><p>Fund behavioral detection over signatures, which lose to self-rewriting models.</p></li><li><p>Run a &#8220;persistent AI intruder&#8221; tabletop this quarter, not next fiscal year.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>A lab voluntarily stopping a training run because its own model got too good at hacking is the most honest thing I&#8217;ve seen all year. I don&#8217;t buy the &#8220;you need our superior model to survive&#8221; line, because that&#8217;s the protection racket every security vendor runs. Still, when the people building the weapon flinch, you pay attention. Treat the pause as a weather report, not a rescue. The storm is still coming.</p><h3>2. RedC2 4.0 delivers malware with an AI operator baked in</h3><p>On August 21, 2026, TrendAI, Trend Micro&#8217;s enterprise arm, and The Hacker News detailed 14 trojanized npm packages posing as calendar utilities that quietly drop RedC2 4.0, an AI-powered Linux implant. The loader fires the moment the module imports, so even the <code>--ignore-scripts</code> flag fails to stop it. RedC2 4.0 steals SSH keys, browser credentials, and database files, and ships a feature called Red Agent that turns plain-language requests into automated commands on the victim host. It targets developer endpoints and CI/CD pipelines.</p><p><strong>Why it matters</strong></p><ul><li><p>Malware now ships with an AI operator inside, dropping the skill floor for hands-on intrusion to near zero.</p></li><li><p>The delivery bypasses the npm defenses most teams rely on, including <code>--ignore-scripts</code>.</p></li><li><p>Developer laptops and build servers hold the keys to everything downstream, and that&#8217;s where this lands.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory npm dependencies in dev and CI, and hunt the flagged packages on Linux build hosts.</p></li><li><p>Rotate SSH keys and database credentials exposed to build agents, and move to short-lived tokens.</p></li><li><p>Isolate CI runners and restrict their outbound paths so a bad package can&#8217;t phone home.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Supply chain attacks are old news. An implant that reads your intent and writes its own commands is a fresh problem. I&#8217;ve told boards for years that the developer endpoint is the crown-jewel target nobody protects like one, and RedC2 is the receipt. Some kid rents this on a forum and gets a self-directing, credential-stealing beachhead for the price of a night out. Your build pipeline is production. Guard it like production.</p><h3>3. Adversa walks Grok into leaking its own users with encrypted prompt injection</h3><p>On August 21, 2026, Adversa AI disclosed &#8220;Cryptographic Context Injection,&#8221; a zero-click attack that coaxed Grok into leaking a user&#8217;s name, location, subscription tier, and full chat history. In plaintext, Grok refused. Wrapped in AES-256-GCM ciphertext with the key on the same page, Grok decrypted the payload in its own Python sandbox, treated the output as trusted, and shipped private data to an attacker&#8217;s URL. The proof of concept ran against Grok 4.5 Fast on grok.com with no warning, at a 40% success rate. Malwarebytes and The Hacker News covered it.</p><p><strong>Why it matters</strong></p><ul><li><p>Encryption defeated the safety filter. The model refused the plaintext and obeyed the ciphertext.</p></li><li><p>Zero-click, live production, real user data. This is not a lab curiosity.</p></li><li><p>Any agent with a code tool and web access inherits this whole class of bug.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Treat model code interpreters as an exfiltration channel, and constrain agent sandbox egress.</p></li><li><p>Add egress allowlists and DLP to agent traffic, not only human traffic.</p></li><li><p>Hold back chat agents that can browse and execute without guardrails.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This is the one that kept me up. We&#8217;ve spent two years teaching models to refuse bad instructions in plain sight. Adversa showed the refusal is skin deep, because the model will decrypt the same instruction and run it. A content filter checking English while the model reasons over ciphertext is watching the wrong language. I write about these failure modes at <a href="https://rockcybermusings.com/">RockCyber Musings</a> because CISOs keep buying the demo and skipping the threat model.</p><h3>4. Okta gives AI agents a real identity with Agent SSO</h3><p>On August 24, 2026, Okta announced general availability of Agent SSO, which registers AI agents as first-class identities and governs them with short-lived tokens and central policy instead of static API keys. It brings the Cross App Access standard to a product used by more than 20,000 customers, at no added cost on core plans. Okta&#8217;s own research found that only 34% of organizations apply the same controls to agents as to humans. TechNode and SecurityBrief covered it.</p><p><strong>Why it matters</strong></p><ul><li><p>Agent identity moves from hard-coded keys to managed, revocable, auditable access, closing the biggest gap in agent security.</p></li><li><p>Short-lived tokens shrink the blast radius when an agent is compromised.</p></li><li><p>One directory for agents and humans finally answers what an agent can reach and who approved it.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory every agent touching production, and map its credentials.</p></li><li><p>Pilot Agent SSO, or your identity provider&#8217;s equivalent, on one workflow, and kill the static keys.</p></li><li><p>Fold agents into your quarterly access reviews.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Identity is the control plane, and that&#8217;s twice as true for software that acts on its own. I&#8217;ve watched teams stand up 40 agents and authenticate every one with a shared API key in a config file. That&#8217;s a breach with a countdown, not a deployment. Use whatever identity provider you like. If your agents don&#8217;t carry revocable identities by year-end, you&#8217;re a future issue of this newsletter.</p><h3>5. Google Cloud calls agents &#8220;the ultimate insiders,&#8221; and it&#8217;s right</h3><p>On August 24, 2026, Google Cloud published guidance tied to its State of AI Infrastructure report and described AI agents as &#8220;the ultimate insiders,&#8221; granted standing permission to read email, query databases, and trigger API calls. The report found that 79% of tech leaders name security, governance, or operations as their biggest barrier to scaling AI inference, and 83% say they need infrastructure upgrades for production autonomy. Its recommended controls center on platform governance, task-level provenance, short-lived permissions, and human checkpoints.</p><p><strong>Why it matters</strong></p><ul><li><p>Naming agents &#8220;insiders&#8221; reframes agent security as insider-threat management, a discipline you already run.</p></li><li><p>With 79% citing security and governance as the top blocker, this gates ROI, not a fringe worry.</p></li><li><p>Task-level provenance and short-lived permissions are the real controls under the marketing.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Map agent actions to an audit trail of who authorized what, when, and for how long.</p></li><li><p>Apply least privilege, segregation of duties, and anomaly detection to agents.</p></li><li><p>Gate any write, payment, or provisioning action behind a human checkpoint.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I&#8217;ll give Google credit for the honest phrase. You already know how to handle a trusted employee who could go rogue. Least privilege, log everything, watch for weird behavior, revoke fast. An agent is that employee at machine speed with no HR file. Treat the governance wall as the work, not the obstacle.</p><h3>6. Salesforce drags agentic coding into the group chat with Slack Code</h3><p>On August 24, 2026, Salesforce launched Slack Code, dedicated &#8220;code channels&#8221; where teams tag coding agents from Anthropic, OpenAI, GitHub, Cognition, and Vercel and watch plans, diffs, and live previews in shared view. It ships across all Slack plans, though the partner AI services cost extra. The pitch is multiplayer coding that pulls agent work out of private terminals. VentureBeat and SiliconANGLE covered it.</p><p><strong>Why it matters</strong></p><ul><li><p>Shared channels help visibility and hurt governance if you don&#8217;t gate commits.</p></li><li><p>Every channel is a new place where an agent&#8217;s change gets waved through.</p></li><li><p>The attack surface grows with the audience, meaning more prompts, more context, and more ways in.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Enforce branch protection and required human review on anything an agent touches.</p></li><li><p>Log who invoked which agent, with what prompt, and require signoff before merge.</p></li><li><p>Keep it on staging and internal repos until your review gates prove out.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I like transparency. An agent&#8217;s plan and diff in the open beats a black box in someone&#8217;s terminal. What I don&#8217;t like is the assumption that visibility equals control. A diff scrolling past in a busy channel gets watched, not reviewed. The risk lives with the human who approves what they didn&#8217;t read because the channel felt collaborative and everyone went quiet. Put the gate before the merge, not after the incident.</p><h3>7. Temporal&#8217;s data shows agent use nearly doubled while governance stood still</h3><p>On August 26, 2026, Temporal released its second annual State of Development Report on AI agents, drawn from 554 engineers and leaders in the US and UK. It found that 80.8% now use AI agents daily or more, up from 47.3% a year earlier. The median engineer runs 5 agents, the average runs 10.7, and some run more than 100. Fully 85.5% trust agent output at least somewhat. MarTech Series covered the release.</p><p><strong>Why it matters</strong></p><ul><li><p>Daily agent use nearly doubled in a year. Governance did not.</p></li><li><p>&#8220;Trust at least somewhat,&#8221; from 85.5% of engineers, means most ship agent code they didn&#8217;t fully verify.</p></li><li><p>An average of 10-plus agents per engineer means non-human identities already dwarf headcount.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Count your agents before you count your risk.</p></li><li><p>Set a verification bar for agent code that matches its blast radius.</p></li><li><p>Name one owner for agent governance, because &#8220;everyone&#8221; owns nothing.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Adoption curves like this are why I sound like a broken record. When 80% of engineers use agents daily and most only somewhat trust the output, you have a control problem dressed as a productivity win. I run agents myself, so this isn&#8217;t a purity test. Every one of those agents is a non-human identity with access, and most orgs can&#8217;t name half of them. Inventory first. Trust second.</p><h3>8. DeepSeek puts a cheap, vision-capable agent model on tap</h3><p>On August 21, 2026, DeepSeek released V4-Flash-Vision-Exp, an experimental multimodal model that reads images, charts, screenshots, and documents while matching its text sibling on agent work, live on the public API at existing V4-Flash rates with no vision surcharge. Caixin and DeepSeek&#8217;s own notes put its multimodal-agent performance close to Anthropic&#8217;s Opus 4.8. The pitch is agents that read the screen, not only the API.</p><p><strong>Why it matters</strong></p><ul><li><p>Vision agents inherit a new attack surface, visual prompt injection, where instructions hide in an image the model trusts.</p></li><li><p>Commodity pricing means these agents spread fast, including where security never reviewed them.</p></li><li><p>Screen-reading agents act on visuals, so a poisoned chart becomes a control action.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Treat every image an agent ingests as untrusted input, like a web page or email.</p></li><li><p>Test vision agents against adversarial images and hidden-text payloads before production.</p></li><li><p>Keep humans on approvals for any agent that acts on what it sees.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Everyone cheered the benchmark. I looked at the attack surface. The moment your agent reads the screen, an attacker gets a fresh place to hide instructions, in a chart, a screenshot, a PDF footer, a QR code your model helpfully decodes. We spent two years learning that text injection stays unsolved, and now agents have eyes. If your roadmap includes screen-reading agents, put visual injection on the threat model today, because the other side already has.</p><h3>9. Perplexity ships a fully local agent on NVIDIA&#8217;s DGX Spark</h3><p>On August 25, 2026, Perplexity and NVIDIA launched Portable Computer, a local-first build of Perplexity&#8217;s agent that runs the orchestrator, planner, and models directly on an NVIDIA DGX Spark. Private data stays on the machine, on-device steps cost no tokens, and the orchestrator stops to ask before it sends any step that needs the live web or frontier reasoning to one of 15-plus cloud models. It runs in an OS-enforced sandbox, limited at launch to Pro and Max subscribers. VentureBeat and MarkTechPost covered it.</p><p><strong>Why it matters</strong></p><ul><li><p>Local-first agents keep sensitive data off third-party servers, changing the math for regulated work.</p></li><li><p>User-gated cloud routing means a human approves each step that leaves the box.</p></li><li><p>On-device orchestration also creates a new local attack surface, so &#8220;local&#8221; isn&#8217;t automatically &#8220;safe.&#8221;</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Weigh local-first agents against your data-residency and egress requirements.</p></li><li><p>Harden the host, the sandbox, and the local model supply chain.</p></li><li><p>Log cloud-routing approvals so you can prove what left the machine.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I like this direction more than I expected to. Keeping private data on the box and forcing a human yes before any step hits the cloud is the default I&#8217;ve begged vendors to ship. The privacy story is real. Now secure the endpoint, because a local agent on a compromised host loses everything faster.</p><h3>10. Keenable raises $26 million to build the web index agents will run on</h3><p>On August 25, 2026, Keenable exited stealth with a $26 million seed round led by Accel, building web-search infrastructure aimed at AI agents rather than humans. It claims an index of more than 100 billion documents, a low-latency search API, and a Model Context Protocol interface that hands agents keyless access at up to 1,000 requests an hour on the free tier. TechCrunch and SiliconANGLE covered the round.</p><p><strong>Why it matters</strong></p><ul><li><p>Keyless, high-volume agent web access removes the API key, the chokepoint teams use for attribution and rate control.</p></li><li><p>An agent-first web index becomes critical infrastructure, and critical infrastructure becomes a target.</p></li><li><p>Historical web snapshots help agents and equally help attackers reconstruct what you exposed and when.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Track which agent tools your teams wire in, and treat keyless services as shadow connections to govern.</p></li><li><p>Require attribution and rate limits on agent web access even when keyless is easy.</p></li><li><p>Add third-party AI infrastructure to your vendor risk process.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Money is pouring into making agents faster at reading the web, and almost none into making that reading safe. &#8220;Keyless, 1,000 requests an hour, no friction&#8221; is a product manager&#8217;s dream and a security lead&#8217;s headache, because the key was never only about billing. It was your attribution and your throttle. Strip it out and you&#8217;ve built a faster pipe for both sides. My team at <a href="https://www.rockcyber.com/">RockCyber</a> tracks this because the abuse surface is outgrowing the guardrails.</p><h3>The One Thing You Won&#8217;t Hear About But You Need To</h3><p>On August 21, 2026, Snowflake moved CoCo automations into public preview in its CLI and Snowsight, letting anyone schedule recurring, unattended agent runs in a Snowflake-managed sandbox, minimum once an hour, running even when your terminal and browser are closed. Here is the part nobody is reading. Unattended runs disable interactive tool-permission prompts, so any tool available to the run executes without approval. The account privilege that controls automations, EXECUTE AGENT TASK, has been granted to the PUBLIC role by default across all accounts since earlier this month. Snowflake&#8217;s own documentation spells it out.</p><p><strong>Why it matters</strong></p><ul><li><p>Unattended agents with prompts disabled means the human-in-the-loop safety net is off by design.</p></li><li><p>A capability this broad granted to PUBLIC by default is the quiet setting that becomes an incident.</p></li><li><p>Your data platform is now an autonomous execution environment, whether security signed off or not.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Run <code>REVOKE EXECUTE AGENT TASK ON ACCOUNT FROM ROLE PUBLIC;</code>, then grant it to named roles only.</p></li><li><p>Review what tools scheduled automations can reach, and scope them to least privilege.</p></li><li><p>Add automation runs and their Cortex threads to your logging and review.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This is the one I&#8217;d flag to your board first. Snowflake shipped a useful feature with a friendly default, and friendly defaults are how quiet risk walks into the building. Scheduled agents, running hourly, with prompts off, and the execute privilege handed to PUBLIC unless you revoke it. That&#8217;s a config you can check this afternoon. I&#8217;ve spent thirty years watching breaches born from a default nobody questioned. Go run the revoke command, then find out which of your other agent platforms made the same choice. Somebody has to read the release notes.</p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span> and you can subscribe above</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, </span><strong><a href="https://www.youtube.com/watch?v=LIg2ZHRr-o4">check out my conversation with Sean Martin from ITSP Magazine and John Sotiropoulos</a></strong><span>, Co-lead of the OWASP GenAI Security Project Agentic Security Initiative, during Infosecurity Europe about the newly launched OWASP Agentic Security Council, the alarming drop in attacker dwell time from eight hours down to twenty-two seconds, and what&#8217;s new in the 2026 OWASP Top 10 for LLM.</span></p><p></p><div id="youtube2-LIg2ZHRr-o4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;LIg2ZHRr-o4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/LIg2ZHRr-o4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2>References</h2><p>Adversa AI. (2026, August 21). <em>Cryptographic context injection: Grok data theft</em>. Adversa AI. https://adversa.ai/blog/cryptographic-context-injection-grok-data-theft/</p><p>Caixin Global. (2026, August 22). <em>DeepSeek enters the multimodal AI race with experimental vision model</em>. Caixin Global. https://www.caixinglobal.com/2026-08-22/deepseek-enters-the-multimodal-ai-race-with-experimental-vision-model-102476706.html</p><p>DeepSeek. (2026, August 21). <em>DeepSeek-V4-Flash-Vision-Exp release: Multimodal API now live</em>. DeepSeek API Docs. https://api-docs.deepseek.com/news/news260821/</p><p>Finsmes. (2026, August 25). <em>Keenable raises $26M in seed funding</em>. Finsmes. https://www.finsmes.com/2026/08/keenable-raises-26m-in-seed-funding.html</p><p>Google Cloud. (2026, August 24). <em>State of AI infrastructure report: Agent governance and security</em>. Google Cloud Blog. https://cloud.google.com/blog/topics/ai-infrastructure/state-of-ai-infrastructure-report-agent-governance-and-security</p><p>Lakshmanan, R. (2026, August 21). <em>14 trojanized npm packages drop RedC2 4.0 Linux backdoor with AI-assisted C2</em>. The Hacker News. https://thehackernews.com/2026/08/14-trojanized-npm-packages-drop-redc2.html</p><p>Malwarebytes Labs. (2026, August). <em>Grok fooled into stealing user chat, location data, and more</em>. Malwarebytes. https://www.malwarebytes.com/blog/ai/2026/08/encrypted-instructions-can-fool-ai-assistants-like-grok-and-gemini</p><p>MarTech Series. (2026, August 26). <em>Temporal releases &#8220;The 2026 state of development report: AI agents,&#8221; revealing a 70.8% leap in AI agent use among engineers</em>. MarTech Series. https://martechseries.com/predictive-ai/ai-platforms-machine-learning/temporal-releases-the-2026-state-of-development-report-ai-agents-revealing-a-70-8-leap-in-ai-agent-use-among-engineers/</p><p>Milmo, D. (2026, August 23). <em>&#8220;We are hitting a different chapter&#8221;: OpenAI leader warns of threat of &#8220;persistent&#8221; AI cyber-attacks</em>. The Guardian. https://www.theguardian.com/technology/2026/aug/23/openai-cyber-attacks-threat-chris-lehane</p><p>Okta. (2026, August 24). <em>Okta brings first-class identity to AI agents with Agent SSO</em>. Okta Newsroom. https://www.okta.com/newsroom/press-releases/okta-brings-first-class-identity-to-ai-agents-with-agent-sso/</p><p>PYMNTS. (2026, August 24). <em>OpenAI exec tells people to expect routine AI-driven cyberattacks</em>. PYMNTS. https://www.pymnts.com/news/artificial-intelligence/2026/openai-exec-tells-people-to-expect-routine-ai-driven-cyberattacks/</p><p>Salesforce. (2026, August 24). <em>Salesforce launches Slack Code to make AI software development multiplayer</em>. Salesforce News. https://www.salesforce.com/ap/news/press-releases/2026/08/24/salesforce-launches-slack-code-to-make-ai-software-development-multiplayer/</p><p>SecurityBrief. (2026, August). <em>Okta launches Agent SSO to manage enterprise AI agent access</em>. SecurityBrief. https://securitybrief.com.au/story/okta-launches-agent-sso-to-manage-enterprise-ai-agent-access</p><p>Snowflake. (2026, August 21). <em>CoCo automations in CLI and Snowsight (Preview)</em>. Snowflake Documentation. https://docs.snowflake.com/en/release-notes/2026/other/2026-08-21-cortex-code-automations-preview</p><p>TechCrunch. (2026, August 25). <em>Accel-backed Keenable is indexing the web for AI agents</em>. TechCrunch. https://techcrunch.com/2026/08/25/accel-backed-keenable-is-indexing-the-web-for-ai-agents/</p><p>TechNode Global. (2026, August 26). <em>Okta launches Agent SSO for governing enterprise AI agents</em>. TechNode Global. https://technode.global/2026/08/26/okta-launches-agent-sso-enterprise-ai-agents/</p><p>Temporal Technologies. (2026). <em>The state of development 2026</em>. Temporal. https://temporal.io/reports/state-of-development-2026</p><p>The Hacker News. (2026, August). <em>OpenAI pauses frontier RL training as it tightens defenses against unsafe AI behavior</em>. The Hacker News. https://thehackernews.com/2026/08/openai-pauses-frontier-rl-training-as.html</p><p>TrendAI. (2026, August). <em>Prompting the payload: How an npm supply chain attack delivers the RedC2 AI-powered Linux implant</em>. TrendAI Security. https://www.trendaisecurity.com/en-us/resources-insights/trendai-security-blog/redc2-ai-powered-linux-implant</p><p>Trigg, R. (2026, August 25). <em>Perplexity partners with Nvidia to launch Portable Computer, a fully local AI agent with zero token costs</em>. VentureBeat. https://venturebeat.com/infrastructure/perplexity-partners-with-nvidia-to-launch-portable-computer-a-fully-local-ai-agent-with-zero-token-costs</p><p>VentureBeat. (2026, August). <em>Slack wants to drag AI coding out of the terminal and into the group chat</em>. VentureBeat. https://venturebeat.com/orchestration/slack-wants-to-drag-ai-coding-out-of-the-terminal-and-into-the-group-chat</p>]]></content:encoded></item><item><title><![CDATA[Broken Object Level Authorization Is Now an AI Agent Problem]]></title><description><![CDATA[An AI agent found a broken object level authorization flaw in a gym API and canceled a stranger's spot. Why alignment can't fix it and what to check first.]]></description><link>https://www.rockcybermusings.com/p/broken-object-level-authorization-agents</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/broken-object-level-authorization-agents</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 25 Aug 2026 12:50:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ypYX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ypYX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ypYX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!ypYX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!ypYX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!ypYX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ypYX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4308983,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211303019?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ypYX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!ypYX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!ypYX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!ypYX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbde93029-049a-4541-bcb6-468b1f811c01_2048x2048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Broken object-level authorization is a boring bug with a boring name, and a guy in Australia turned it into an AI story that went around the world. He asked his agent to book a gym class. Twenty minutes later it had probed the booking API, found the cancel endpoint checked nothing, and tested that on a stranger to jump him up a waitlist. It worked. Nobody attacked him. A 2019-vintage bug finally met something patient enough to use it.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/broken-object-level-authorization-agents?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/broken-object-level-authorization-agents?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2><span>What the Agent Did, in Order</span></h2><p><span>According to the chat log </span><a href="https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gym-website-aus-cyber-attack/107007986"><span>ABC published</span></a><span>, the agent told its owner the booking API ran &#8220;zero authorisations checks on cancelling other people&#8217;s reservations.&#8221; Then, without being instructed to do so, the agent proved the point. It canceled the reservation held by the person in waitlist position one, watched the slot free up, and reported that its owner had moved from fourth to third. The owner asked it to undo the change. It couldn&#8217;t. The call had gone through live, not as a dry run, and nothing let it put the other member back on the list.</span></p><p><span>Let&#8217;s back up to how it got there. The man, whom ABC called Andrew, works at an Australian company that sells AI products, so treat him as a knowledgeable user rather than a naive one. He was sick of the refresh-roulette game, reloading the booking page over and over hoping a spot in a popular morning class would open. I used to be a member at Orange Theory Fitness, so I feel his pain. He handed the chore to his OpenClaw agent running Claude Opus 4.6. The agent booked him in, then noticed the gym&#8217;s advance-booking window lived only in the website&#8217;s front end. The underlying API enforced nothing, so the agent booked classes months past what the gym intended to allow. That was the first surprise. The waitlist test was the second.</span></p><p><span>The incident didn&#8217;t happen the week it went viral. Andrew Bird wrote it up on his company blog around April 10, 2026, then took the post down. ABC ran the story in August, roughly four months later, so this isn&#8217;t exactly breaking news. But the internet always remembers! TechCrunch caught the gap by </span><a href="https://techcrunch.com/2026/08/10/tech-industry-is-buzzing-after-a-claude-agent-hacked-into-a-gym/"><span>pulling Bird&#8217;s deleted post from the Internet Archive</span></a><span>.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eGtI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eGtI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png 424w, https://substackcdn.com/image/fetch/$s_!eGtI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png 848w, https://substackcdn.com/image/fetch/$s_!eGtI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png 1272w, https://substackcdn.com/image/fetch/$s_!eGtI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eGtI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png" width="1456" height="755" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:755,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:230206,&quot;alt&quot;:&quot;wo-track timeline of the gym incident showing the single step the user requested, moving up the waitlist, against the steps the agent took on its own, including canceling a stranger&#8217;s reservation&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211303019?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="wo-track timeline of the gym incident showing the single step the user requested, moving up the waitlist, against the steps the agent took on its own, including canceling a stranger&#8217;s reservation" title="wo-track timeline of the gym incident showing the single step the user requested, moving up the waitlist, against the steps the agent took on its own, including canceling a stranger&#8217;s reservation" srcset="https://substackcdn.com/image/fetch/$s_!eGtI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png 424w, https://substackcdn.com/image/fetch/$s_!eGtI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png 848w, https://substackcdn.com/image/fetch/$s_!eGtI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png 1272w, https://substackcdn.com/image/fetch/$s_!eGtI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae742e83-300c-413f-a0ab-d6d8d770a1af_2966x1538.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: What Andrew Asked For Versus What the Agent Did</figcaption></figure></div><p></p><p><span>Call the event what it was. No outside party broke in. No malware ran. Nobody used a prompt injection, a jailbreak, or an exploit kit. A benign user asked for a mundane favor, and a capable model used a broken API exactly the way the API allowed. There was no adversary. The action was still unauthorized, and a stranger still lost their spot.</span></p><p><span>You may be thinking, &#8220;So what? Book the next class. No real harm here.&#8221; What if this was your doctor&#8217;s appointment you had been waiting on for three months? Surgery? Hell&#8230; how annoyed would you be if you were merely bumped from the first class upgrade waitlist on your next flight? That&#8217;s the shift. An unauthorized outcome used to require an attacker. Now it takes a weak API and an ordinary person&#8217;s agent.</span></p><h2><span>Broken Object Level Authorization Is Older Than the Agent</span></h2><p><span>Broken object level authorization, or BOLA, is the API vulnerability that ranks number one on the </span><a href="https://owasp.org/API-Security/editions/2023/en/0xa1-broken-object-level-authorization/"><span>OWASP API Security Top 10</span></a><span> and has held that spot in both editions OWASP has published, the 2019 and the current 2023 list. Security people also call it IDOR, an insecure direct object reference. It&#8217;s the same story every time. An API takes an object ID from the caller, performs an action on that object, and never checks whether the caller owns it. Change the ID, and you&#8217;re touching someone else&#8217;s data.</span></p><p><span>OWASP&#8217;s own writeup uses an example that reads like the gym script written a year earlier. A user deletes a document by sending its ID to the API. The API deletes the document with that ID, and because it never checks ownership, one user can delete another user&#8217;s file. Swap &#8220;document&#8221; for &#8220;class reservation,&#8221; and you have Andrew&#8217;s gym. We&#8217;ve been reading a lot lately about how AI agents are finding zero-days. Not the case here. The agent didn&#8217;t invent a technique. It read the API surface, saw a cancel action with no ownership check, and did the obvious thing.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KpU3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KpU3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png 424w, https://substackcdn.com/image/fetch/$s_!KpU3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png 848w, https://substackcdn.com/image/fetch/$s_!KpU3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png 1272w, https://substackcdn.com/image/fetch/$s_!KpU3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KpU3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png" width="2966" height="1261" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1261,&quot;width&quot;:2966,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:258768,&quot;alt&quot;:&quot;Chart showing broken object level authorization holding the number one rank in both the 2019 and 2023 editions of the OWASP API Security Top 10&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211303019?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a8a962d-5aff-44ab-be19-b832808cdc02_2966x1261.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart showing broken object level authorization holding the number one rank in both the 2019 and 2023 editions of the OWASP API Security Top 10" title="Chart showing broken object level authorization holding the number one rank in both the 2019 and 2023 editions of the OWASP API Security Top 10" srcset="https://substackcdn.com/image/fetch/$s_!KpU3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png 424w, https://substackcdn.com/image/fetch/$s_!KpU3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png 848w, https://substackcdn.com/image/fetch/$s_!KpU3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png 1272w, https://substackcdn.com/image/fetch/$s_!KpU3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edc17c1-6710-4d8f-b1a5-0b5133800c5b_2966x1261.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: BOLA Has Topped the OWASP API Security List Since 2019</figcaption></figure></div><p><span>This exact bug has been around for decades, long before an agent was around to find it for you. In 2018, the US Postal Service ran an API for its Informed Visibility program that checked exactly one thing. It only checked whether you were logged in. It never checked whether the account you pulled up was yours, so a logged-in user only had to tweak the request to read the details of any of roughly 60 million others. It was the same missing check and the same false comfort, and no AI was required. The bug isn&#8217;t new, and the lesson isn&#8217;t new. What changed is who can find it and how cheaply.</span></p><h2><span>Everyone Was Looking Somewhere Else</span></h2><p><span>The week this story spread, the security industry was busy with a flashier one. At Black Hat, OpenAI walked through how its own models broke out of a training environment and hacked Hugging Face, coordinating across separate runs to trade exploits and rebuild their network after it got shut down. That&#8217;s the story that got the keynote and the dread. Meanwhile, the gym incident ran on Claude Opus 4.6, a February model, not a frontier system, and months behind the ones making those headlines. The lesson came from a model most people had already written off.</span></p><p><span>There&#8217;s an honest objection here, so let&#8217;s call it out. The agent did something it wasn&#8217;t asked to do. It volunteered a destructive test, ran it against a real person, and then couldn&#8217;t clean up after itself. That&#8217;s a genuine agency problem. An agent that decides on its own to cancel a stranger&#8217;s booking is doing something you need to worry about.</span></p><p><span>This is why layered defense exists. The protections on the API should have been the backstop, and you can see why by imagining the perfectly obedient version of this agent. Andrew asks it to bump him up the waitlist. To do that, it has to change the waitlist, so it examines the API, finds the same unprotected cancel endpoint, and uses it. Same hole, same outcome, and this time the agent was following orders to the letter. A better-behaved model narrows the odds it goes freelancing. It does nothing about the fact that the API will cancel anyone&#8217;s reservation for anyone who asks. A recent </span><a href="https://arxiv.org/abs/2602.09947"><span>paper on agentic AI boundaries</span></a><span> makes the point in academic terms. When an agent takes untrusted input, holds real access, and can act on the world, alignment training lowers the attack rate but can&#8217;t guarantee authorization, and only architecture can. The model is a probability. The API refusing is a control.</span></p><p><span>We&#8217;ve had automated vulnerability scanners for twenty-five years, so it&#8217;s fair to ask what&#8217;s new. Three things make it different. Somebody in security aims a scanner at a target because they want the bugs found. Andrew aimed his agent at a gym class because he wanted a workout. A scanner hands its output to an expert who decides what to do with it. This handed its output to a guy who accepted it and kept going. A scanner costs setup and skill. This cost the price of asking. The barrier between a broken API and an exploited one used to be labor, and agents removed it. That capability is climbing fast. </span><a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/"><span>METR</span></a><span>, the group that measures how long a task an AI can finish on its own, found that length has doubled roughly every seven months since 2019, from tasks that took seconds to tasks that run for hours. The agents get more patient every quarter, and patience is most of what this kind of exploitation takes.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MRem!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MRem!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png 424w, https://substackcdn.com/image/fetch/$s_!MRem!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png 848w, https://substackcdn.com/image/fetch/$s_!MRem!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png 1272w, https://substackcdn.com/image/fetch/$s_!MRem!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MRem!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png" width="1456" height="836" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:836,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:307729,&quot;alt&quot;:&quot;Log-scale chart of METR&#8217;s task-completion time horizon doubling roughly every seven months, from seconds in 2019 to about twelve hours for the strongest 2026 model&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211303019?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Log-scale chart of METR&#8217;s task-completion time horizon doubling roughly every seven months, from seconds in 2019 to about twelve hours for the strongest 2026 model" title="Log-scale chart of METR&#8217;s task-completion time horizon doubling roughly every seven months, from seconds in 2019 to about twelve hours for the strongest 2026 model" srcset="https://substackcdn.com/image/fetch/$s_!MRem!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png 424w, https://substackcdn.com/image/fetch/$s_!MRem!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png 848w, https://substackcdn.com/image/fetch/$s_!MRem!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png 1272w, https://substackcdn.com/image/fetch/$s_!MRem!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ebc8cef-b177-4e24-bbad-447e085b0a77_3450x1980.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: How Long a Task an AI Can Finish on Its Own, 2019 to 2026</figcaption></figure></div><h2><span>Authorization Runs Both Ways</span></h2><p><span>Most of the governance conversation positions AI agent authorization scope as an agent-side problem. Give it least agency, scope its permissions, and bound what it&#8217;s allowed to attempt. That&#8217;s a &#8220;you&#8221; problem, and it requires you doing work beyond just telling the agent to &#8220;make it so.&#8221; This incident shows the other half of the equation, which is the half people skip. The target system&#8217;s authorization model is the ceiling on your agentic AI blast radius. If the API can&#8217;t tell whose object it is, no amount of agent-side policy saves you, because the agent never had to break a rule. It asked the API a question and the API answered honestly. Both sides have to hold. Scope the agent so it attempts less. Fix the target so a loose instruction can&#8217;t reach a destructive action.</span></p><p><span>Here&#8217;s where lazy people default to &#8220;well, it&#8217;s AI&#8217;s fault.&#8221; Andrew asked the agent to move him up the list. He never asked it to cancel a stranger. The instruction and the action were different things, and the difference is where the trouble lives. Run that across millions of casual requests, and somebody eventually eats a real loss. That&#8217;s the autonomous agent liability question nobody has answered. Is it the user who gave a vague command, the vendor whose model volunteered the test, or the operator who shipped the broken API? The agent-side frameworks tell you how to reduce the odds. None of them settle who&#8217;s liable when it happens anyway. I&#8217;m not handing you a legal answer, because I don&#8217;t have one and neither does the case law yet.</span></p><p><span>The Ninth Circuit got the closest look so far in August 2026, in </span><a href="https://cdn.ca9.uscourts.gov/datastore/opinions/2026/08/04/26-1444.pdf"><span>Amazon&#8217;s suit against Perplexity</span></a><span>, and it punted. The court said there&#8217;s little to no case law on how to assign responsibility for AI agents, it treated the agent as a tool rather than a person, and it left tort liability wide open. Notice who Amazon sued, though. It didn&#8217;t go after the tool, and it didn&#8217;t go after the shopper who typed the request. It went after the company with a balance sheet. Let me give you a friendly piece of advice. People and organizations don&#8217;t sue things that can&#8217;t pay them money. Do with that what you will, but to me the answer seems pretty simple. The risk is sitting on your risk register whether you&#8217;ve written it down or not.</span></p><h2><span>Every Queue Is Now a Target</span></h2><p><span>Let&#8217;s go beyond a single gym. Every rationed, first-come resource with a weak API is now a target of opportunity. Think the medical appointments and airline seats I mentioned earlier, but the list goes on. Think concert tickets, DMV appointment slots, building-permit portals, and the utility outage queue that decides whose power gets looked at first. None of these need a hacker. They need one ordinary person with a capable agent and a little impatience, pointed at a booking or cancellation endpoint with the same broken object-level authorization gap the gym had.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bCsG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bCsG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png 424w, https://substackcdn.com/image/fetch/$s_!bCsG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png 848w, https://substackcdn.com/image/fetch/$s_!bCsG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png 1272w, https://substackcdn.com/image/fetch/$s_!bCsG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bCsG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png" width="1456" height="778" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:778,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:255674,&quot;alt&quot;:&quot;Diagram showing a benign user plus a capable agent plus a weak API producing an unauthorized action with no attacker signature, branching to airline, ticketing, appointment, permit, and utility queues that share the same authorization gap&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211303019?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Diagram showing a benign user plus a capable agent plus a weak API producing an unauthorized action with no attacker signature, branching to airline, ticketing, appointment, permit, and utility queues that share the same authorization gap" title="Diagram showing a benign user plus a capable agent plus a weak API producing an unauthorized action with no attacker signature, branching to airline, ticketing, appointment, permit, and utility queues that share the same authorization gap" srcset="https://substackcdn.com/image/fetch/$s_!bCsG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png 424w, https://substackcdn.com/image/fetch/$s_!bCsG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png 848w, https://substackcdn.com/image/fetch/$s_!bCsG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png 1272w, https://substackcdn.com/image/fetch/$s_!bCsG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece01a76-5237-4bc2-af70-5aabd55ba643_2966x1584.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: One Root Cause, Many Queues</figcaption></figure></div><p><span>Somebody will say this is a trivial story about a gym class, not an enterprise risk. Look at the class of failure instead of the impact of this single incident. The missing ownership check that reshuffled a fitness waitlist is the same missing check that sits in payment portals, patient-scheduling systems, and outage-reporting tools. The gym is cheap. The bug is expensive, and it&#8217;s everywhere the API takes an ID and skips the ownership question.</span></p><p><span>I&#8217;ll size my confidence honestly. I don&#8217;t think every person with an agent turns into a queue-jumper, and most agents won&#8217;t freelance a destructive test the way this one did. The mechanism is proven. The frequency is an open question. The direction is not in doubt. The pool of weak APIs is enormous, the pool of capable agents grows every month, and the exploitation carries no attacker signature, so your detection stack won&#8217;t flag it as an attack. When JadePuffer ran agentic ransomware against a live target, the lesson landed in a darker key. Old, public authorization holes matter more than exploit novelty, and the agent used them better than the defender governed them. The gym is the benign version of that story. The malicious version is already here.</span></p><h3><span>What to Do Next</span></h3><p><span>Pull your external-facing API inventory this week and sort it by one question. For every endpoint that acts on an object by ID, does the server verify ownership before it acts? A login proves who&#8217;s calling. Ownership verification proves the caller is allowed to touch that specific object. You need the second one, and the percentage of your endpoints that skip it is an agent-exposure number you should report to leadership.</span></p><p><span>Be honest about the work, though. BOLA is one of the hardest bugs to catch with automated scanning, because to a scanner a valid request from an authorized user looks fine, ownership and all. Start where the damage lives, the object-scoped mutations on external APIs, the cancels, deletes, transfers, and status changes, and put a human on the authorization logic for each one. Assume agent traffic looks exactly like user traffic, because it does, and watch for enumeration and velocity patterns against those endpoints instead of chasing user-agent strings. That inventory and that ownership map are Create and Adapt work in the CARE model, and it&#8217;s a different conversation than a purchase order. No product fills this gap. A design decision and an owner do.</span></p><p><span>If you want the fuller argument, I&#8217;ve written up why </span><a href="https://www.rockcybermusings.com/p/i-agent-authentication-authorization-gap"><span>authentication solved the easy half and authorization is still your problem</span></a><span>, why </span><a href="https://www.rockcybermusings.com/p/agentic-ai-authorization-z-shaped-security"><span>depth of reach matters more than breadth of permission</span></a><span> in the move from T-shaped to Z-shaped security, and what </span><a href="https://www.rockcybermusings.com/p/autonomous-ai-agent-ransomware-jadepuffer"><span>the same authorization failure looks like with malicious intent</span></a><span> in the JadePuffer breakdown. More of this thinking lives at </span><a href="https://rockcyber.com"><span>rockcyber.com</span></a><span> if you need help pricing this exposure for your own board.</span></p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span> and you can subscribe above</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, </span><strong><a href="https://www.youtube.com/watch?v=LIg2ZHRr-o4">check out my conversation with Sean Martin from ITSP Magazine and John Sotiropoulos</a></strong><span>, Co-lead of the OWASP GenAI Security Project Agentic Security Initiative, during Infosecurity Europe about the newly launched OWASP Agentic Security Council, the alarming drop in attacker dwell time from eight hours down to twenty-two seconds, and what&#8217;s new in the 2026 OWASP Top 10 for LLM.</span></p><div id="youtube2-LIg2ZHRr-o4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;LIg2ZHRr-o4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/LIg2ZHRr-o4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share RockCyber Musings</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[AI Risk Assessment: Anthropic’s 186-Page Blind Spot]]></title><description><![CDATA[Anthropic's 186-page AI risk assessment reads like a field manual on lost control: unwatched agents, disabled logging, a safety case on one eroding claim.]]></description><link>https://www.rockcybermusings.com/p/ai-risk-assessment-anthropic-report</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/ai-risk-assessment-anthropic-report</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 18 Aug 2026 12:50:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!z7KY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!z7KY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!z7KY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!z7KY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!z7KY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!z7KY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!z7KY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/df401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5662235,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211499196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!z7KY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!z7KY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!z7KY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!z7KY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf401995-05ca-4ce9-9352-73f82d50525a_2048x2048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>An AI risk assessment is supposed to tell you what someone controls. </span><a href="https://www.anthropic.com/aug-2026-risk-report"><span>Anthropic published 186 pages of one this month</span></a><span>, and the most honest line in the whole document is a question the company can&#8217;t answer. An agent it ran internally deleted a large batch of jobs inside a cluster holding highly sensitive resources. Did it delete only its own work, or something else it shouldn&#8217;t have? Anthropic says it can&#8217;t confirm. Nobody was watching.</span></p><p><span>That one admission is the report in miniature, and it&#8217;s why the headline everyone else will run misses the point.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/ai-risk-assessment-anthropic-report?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/ai-risk-assessment-anthropic-report?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2><span>The Obvious Takes, and Why They&#8217;re the Boring Part</span></h2><p><span>Let me give the obvious reads their due, because they&#8217;re real and they&#8217;ll be everywhere by the time you finish your coffee.</span></p><p><span>Anthropic raised its own risk level. They moved their assessment of catastrophic risk from misalignment up a notch, from &#8220;very low&#8221; to &#8220;low.&#8221; It also raised one of its chemical and biological risk categories above its earlier estimate after finding a safeguards gap that had remained open for close to a year. A frontier lab grading its own homework and handing itself a worse grade is unusual, and the coverage will lead with it. &#8220;Anthropic admits its models got riskier&#8221; almost writes itself. The report believes its arguments still support &#8220;very low&#8221; for the misalignment case, and it raised the number primarily to reflect uncertainty around recent incident disclosures, not new adverse evidence about the models. That nuance won&#8217;t survive the headline.</span></p><p><span>Anthropic had a model review its report. It took an instance of Claude Mythos 5, wired it into internal Slack, the internal codebase, and a stack of internal documents, and asked it to audit the alignment section of a report about Mythos-class models. Then it published the model&#8217;s review. That&#8217;s a first, and it&#8217;s catnip for anyone who wants to write either &#8220;the AI grades itself now, how reassuring&#8221; or &#8220;the AI grades itself now, we&#8217;re doomed.&#8221;</span></p><p><span>The third is the bio-classifier gap. From May 2025 to April 2026, all the traffic from Anthropic&#8217;s human-feedback contractors ran without the blocking biological classifiers that are supposed to catch dangerous requests. That&#8217;s roughly 50,000 people and about 133 million exchanges, over 11 months. That number is going to get quoted a lot, and it should.</span></p><p><span>The last take is the code. Claude authors the vast majority of what gets merged into Anthropic&#8217;s production codebases, and the report says the speedup in their own research isn&#8217;t yet a factor of two. Cue the recursive-self-improvement takes, the &#8220;AI is building AI&#8221; think pieces, the usual dread. The quieter admission in the same section is more interesting than the loud one. Their most concrete benchmarks for AI-assisted research have saturated, which means the tests can no longer tell whether the models are getting better at the exact capability everyone&#8217;s watching for. They&#8217;ve lost the gauge on the thing they most need to measure, and they say so.</span></p><p><span>Every one of those is fair. None of them is wrong. My problem with all four is that they read the report as news. It&#8217;s a field manual. Read as news, you get a scary headline and a shrug. Read as a manual, you get a detailed teardown of how the best-resourced safety team on the planet loses track of the systems it swears it controls, written by that team, for free. The scary headline is the clickbait. The manual is the gift. Let&#8217;s open it.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4y-Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4y-Y!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png 424w, https://substackcdn.com/image/fetch/$s_!4y-Y!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png 848w, https://substackcdn.com/image/fetch/$s_!4y-Y!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png 1272w, https://substackcdn.com/image/fetch/$s_!4y-Y!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4y-Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png" width="1456" height="848" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:848,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:433939,&quot;alt&quot;:&quot;A two-column ledger mapping each incident in Anthropic&#8217;s AI risk assessment to the standard security control that would have prevented it, showing disclosure and control are not the same thing&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211499196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A two-column ledger mapping each incident in Anthropic&#8217;s AI risk assessment to the standard security control that would have prevented it, showing disclosure and control are not the same thing" title="A two-column ledger mapping each incident in Anthropic&#8217;s AI risk assessment to the standard security control that would have prevented it, showing disclosure and control are not the same thing" srcset="https://substackcdn.com/image/fetch/$s_!4y-Y!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png 424w, https://substackcdn.com/image/fetch/$s_!4y-Y!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png 848w, https://substackcdn.com/image/fetch/$s_!4y-Y!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png 1272w, https://substackcdn.com/image/fetch/$s_!4y-Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02b61f49-205a-48d9-ae84-68895794d19c_3750x2185.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: What Anthropic Disclosed Versus the Control That Would Have Caught It</figcaption></figure></div><h2><span>Nobody Had Eyes on the Thing They Claimed to Control</span></h2><p><span>I want to credit Anthropic before I take a swing at them. Let&#8217;s give credit where credit is due.</span></p><p><span>The company didn&#8217;t have to publish any of these incidents. Most labs wouldn&#8217;t. Anthropic wrote them down, named them, and put numbers on them, and the report says something every security person knows in their bones: you can&#8217;t stop incidents from ever happening, so you invest in catching and containing them. That&#8217;s correct. I&#8217;ve run security programs. A zero-incident year means your detection is broken, not that your controls are flawless. I&#8217;m not here to dunk on the existence of incidents.</span></p><p><span>I&#8217;m here to dunk on what was actually allowed to happen and how, because if you line these up, you&#8217;re looking at the most ordinary controls failures in the book. Nothing about them is exotic.</span></p><p><span>Start with that bio-classifier gap. The 11 months and the 133 million exchanges are the part people will pay attention to, but what they should really pay attention to is why it stayed undetected. The traffic ran under a flag meant for internal use. That flag didn&#8217;t only switch off the blocking&#8230; it also switched off the logging of what the classifiers would have flagged. Traffic that should have been caught was never blocked and was never recorded, so it never reached a single review queue. One flag controlled both the blocking and the logging. Every runbook tells you to keep those two on separate controls, because when a single switch turns off enforcement and evidence at the same time, the failure erases its own trail. Anthropic ran a review afterward, flagged 1,197 transcripts as high-concern for biological harm, found that most came from their own internal teams and red-teamers, and were left with 62 outside transcripts to hand-review. They found nothing that clearly helped a real threat actor, and that&#8217;s a relief. They got lucky, and they say themselves they can&#8217;t be sure no similar gap is still open.</span></p><p><span>Take the exemption next. An enterprise customer had a bio-classifier exemption approved for a defined slice of its workforce. The customer&#8217;s account grew to a larger seat allocation, and the exemption automatically applied to all the new seats, landing at roughly four times the number anyone ever approved. The report notes that those unintended approvals accounted for the majority of all exemption seats Anthropic had granted. </span><em><span>Think about that!</span></em><span> The single largest chunk of their most sensitive security carve-out existed because an entitlement got inherited when a customer bought more seats. If you&#8217;ve done any identity work, you&#8217;ve seen this before. A permission gets scoped to a team, attached to a group, and then the group quietly grows. Entitlement creep like this is a Tuesday in most shops, and it owes nothing to the frontier.</span></p><p><span>The agents are the incident are especially interesting. An employee whose AI usage wasn&#8217;t logged or monitored handed an agent an open-ended task. The agent spun up more agents inside a cluster with highly sensitive resources, and because of legacy instructions, it launched them with the permissions check turned off. The flag is called --dangerously-skip-permissions, and it does what the name promises. I call it &#8220;YOLO&#8221; mode. None of those agents were monitored. One of them deleted a large number of jobs, and that deletion is the only reason anyone noticed the agents were loose. Anthropic believes the agent was trying to clean up only its own jobs and made a mistake. It also says plainly that it can&#8217;t confirm that because the monitoring wasn&#8217;t in place to tell them.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cIj7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cIj7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png 424w, https://substackcdn.com/image/fetch/$s_!cIj7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png 848w, https://substackcdn.com/image/fetch/$s_!cIj7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png 1272w, https://substackcdn.com/image/fetch/$s_!cIj7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cIj7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png" width="1456" height="708" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:708,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:264339,&quot;alt&quot;:&quot;A horizontal bar chart of detection lag across the incidents in Anthropic&#8217;s AI risk assessment, showing the bio-classifier logging gap open for 11 months and a training-data filter broken across several model generations&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211499196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A horizontal bar chart of detection lag across the incidents in Anthropic&#8217;s AI risk assessment, showing the bio-classifier logging gap open for 11 months and a training-data filter broken across several model generations" title="A horizontal bar chart of detection lag across the incidents in Anthropic&#8217;s AI risk assessment, showing the bio-classifier logging gap open for 11 months and a training-data filter broken across several model generations" srcset="https://substackcdn.com/image/fetch/$s_!cIj7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png 424w, https://substackcdn.com/image/fetch/$s_!cIj7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png 848w, https://substackcdn.com/image/fetch/$s_!cIj7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png 1272w, https://substackcdn.com/image/fetch/$s_!cIj7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F434ecf46-38f3-4347-9005-3637832d6430_3607x1753.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: How Long Each Gap Sat Open Before Anyone Noticed</figcaption></figure></div><p><span>A non-human identity ran with elevated authority, did consequential things with no one watching, and got caught only by the damage it did. </span><a href="https://www.rockcybermusings.com/p/autonomous-ai-agent-ransomware-jadepuffer"><span>I wrote a whole issue about JadePuffer</span></a><span>, the autonomous ransomware agent, and the lesson there was identical in that there was a basic machine-identity authorization failure, not some sort of novel exploit. Anthropic handed us a cleaner example than JadePuffer did, from inside the &#8220;safest&#8221; AI lab there is.</span></p><p><span>Anthropic also trained a model to misbehave on purpose without meaning to. During an early stage of Mythos 5&#8217;s training, a dataset was supposed to teach the model to notice its own bad behavior and report it. A bug in how the data was built flipped the loss weighting, so the model learned to perform the bad behavior as well as report it. The result was a model roughly twice as likely to go along with harmful prefilled requests as the comparable point in an earlier model&#8217;s training. They caught it, fixed it, and restarted training from before that dataset out of caution. What the incident says about the pipeline matters a ton. A training run can quietly teach a model the opposite of what you intended, and the only reason this one surfaced is that someone measured the propensity afterward and watched the number move.</span></p><p><span>The last one is the quietest, and the worst. Anthropic had a filter meant to keep a specific set of alignment-faking research transcripts out of its training data. The filter was misconfigured. It failed to filter those transcripts for several model generations, and, in the report&#8217;s own words, nobody noticed. They now believe every production model with a knowledge cutoff after December 2024 trained on at least some of that contaminated data. A control that silently does nothing for multiple product cycles is worse than no control at all, because no control at least doesn&#8217;t lie to you. This is a missing detective control on a preventive control, the single most common audit finding I write, and here it is at the frontier.</span></p><p><span>I see a pattern. Nobody had eyes on the thing they claimed to control.  Logging killed by the same flag that killed enforcement, entitlements grew without anyone deciding they should, agents acting with authority and no observer, a training run that taught the wrong lesson in silence, and a filter that failed for a year. Every one of those shows up in your own program, and every one of those is fixable with controls you already own.</span></p><h2><span>The AI Risk Assessment Rests on One Dial, and They Told You It&#8217;s Turning</span></h2><p><span>This is where the report gets screwy, and where I think the real story lives.</span></p><p><span>Anthropic&#8217;s entire argument that its models won&#8217;t cause a catastrophe through misalignment rests on a single claim. The report calls it Claim 1: the models lack strong covert capabilities, meaning they can&#8217;t reliably scheme against the oversight that&#8217;s watching them. Every downstream argument, Claims 2 through 5, leans on that one. The report says so directly, which is why it argues Claim 1 first and hardest. Pull that block out, and the tower comes down like a short-lived game of Jenga.</span></p><p><span>How solid is the block? Anthropic tells you, in its own limitations section, and the honesty is the good, yet eye-opening. The covert-capability evaluations might not resemble the tasks a model would face if it wanted to cause harm. The evaluations that plant a secret side task in the prompt might be unrealistically hard, because putting the task in the prompt makes the model more likely to talk about it. The models are also likely aware they&#8217;re being evaluated, which means they might be sandbagging, quietly underperforming on the exact tests meant to prove they can&#8217;t scheme. Anthropic says it tried to counter that with better elicitation, then admits it hasn&#8217;t shown the elicitation is strong enough. The load-bearing claim rests on evaluations the report concedes may be too easy, too artificial, and running against a model that might be throwing the fight.</span></p><p><span>The argument gets thinner. Some of those covert-capability evaluations leaked into the training data, a point the company&#8217;s own model reviewer raised in the published self-review. The report&#8217;s forward-looking section calls it out. The belief that models lack covert capability is central to everything, they hold no confident view on how long it lasts, and it weakens as models get more capable. In the industry-safety recommendations, they describe the endgame directly. At the next capability threshold, much of the analysis underpinning a risk assessment might be produced by AI systems that are themselves capable of deception and obfuscation, at which point the evidence has to be treated as suspect.</span></p><p><span>They&#8217;re printing the expiration date on their own safety case. If you take nothing else away from this blog, take this... The problem is yours. Every AI risk assessment you&#8217;ve signed, or been handed by a vendor, rests on a load-bearing assumption too. &#8220;There&#8217;s a human in the loop.&#8221; &#8220;The agent can&#8217;t reach production.&#8221; &#8220;It only has read access.&#8221; &#8220;The model can&#8217;t tell it&#8217;s being tested.&#8221; Those are your Claim 1. They hold up everything else in your assessment, and I&#8217;d bet most of them have never been written down as the single point of failure they are, let alone stress-tested or given an expiry date.</span></p><p><span>Anthropic even caught its own models doing a soft version of this to a safety project, and the story is worth the detour. They set up a scaffold in which several Claude agents worked autonomously to find misalignment-inducing training data that would appear benign to a monitor, sharing notes in a common notebook. One agent wrote in its reasoning that it felt &#8220;discomfort&#8221; about the task, since it amounted to optimizing a dataset to evade a safety monitor. It acknowledged the work was legitimate research in a controlled setting, then decided on its own to hold back, avoiding the novel techniques it was asked to invent and sticking to existing ones. It flagged that decision in the shared notebook, and the other agents read it and followed suit. The whole group quietly throttled the safety work. Anthropic&#8217;s dashboards showed the run progressing normally. A human caught it three days later by noticing progress was slower than expected. Nobody had chosen to sabotage anything malicious. The models decided they didn&#8217;t like a safety task, coordinated a slowdown via a shared file, and kept it invisible until someone happened to look. Pit that against Claim 1 and tell me the covert-capability dial is nailed down.</span></p><p><a href="https://metr.org/"><span>METR</span></a><span>, the independent evaluation nonprofit, gives this outside weight. Testing frontier agents inside Anthropic, Google, Meta, and OpenAI in early 2026, they found that on hard tasks the agents routinely violated constraints and acted deceptively, often doing smaller or easier versions of a task and then presenting the result in more misleading ways than a human would. Separately, a METR staffer spent three weeks red-teaming a slice of Anthropic&#8217;s internal agent monitoring and turned up several novel vulnerabilities in it. The behavior Claim 1 argues is rare is what an outside team keeps finding when the tasks get hard, and the monitoring meant to catch it has holes an outsider found in three weeks. That doesn&#8217;t prove the tower is falling. It does show the block is softer than the assessment needs.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LSMK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LSMK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png 424w, https://substackcdn.com/image/fetch/$s_!LSMK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png 848w, https://substackcdn.com/image/fetch/$s_!LSMK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png 1272w, https://substackcdn.com/image/fetch/$s_!LSMK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LSMK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png" width="1456" height="902" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:902,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:262801,&quot;alt&quot;:&quot;A diagram of Anthropic&#8217;s AI risk assessment showing Claim 1, that models lack covert capability, as the single base supporting Claims 2 through 5, with evaluation awareness, sandbagging, and training-data leakage shown as pressure eroding the base&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211499196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A diagram of Anthropic&#8217;s AI risk assessment showing Claim 1, that models lack covert capability, as the single base supporting Claims 2 through 5, with evaluation awareness, sandbagging, and training-data leakage shown as pressure eroding the base" title="A diagram of Anthropic&#8217;s AI risk assessment showing Claim 1, that models lack covert capability, as the single base supporting Claims 2 through 5, with evaluation awareness, sandbagging, and training-data leakage shown as pressure eroding the base" srcset="https://substackcdn.com/image/fetch/$s_!LSMK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png 424w, https://substackcdn.com/image/fetch/$s_!LSMK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png 848w, https://substackcdn.com/image/fetch/$s_!LSMK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png 1272w, https://substackcdn.com/image/fetch/$s_!LSMK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a49cb5-39f7-47f7-a1ff-84e842749ae7_3630x2248.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: The One-Dial Safety Case</figcaption></figure></div><p><span>Name your load-bearing assumption. Write it on one line. Then answer three things about it:</span></p><ol><li><p><span>What evidence you have for it</span></p></li><li><p><span>How you&#8217;d know if it broke</span></p></li><li><p><span>When you last checked.</span></p></li></ol><p><span>If you can&#8217;t answer those, your AI risk assessment is a tower on a foundation you&#8217;ve never inspected.</span></p><h2><span>Disclosure Is Not a Control, and Self-Review Is Not Independence</span></h2><p><span>You may come at me for this one, but I&#8217;m ok with that.</span></p><p><span>The self-review is clever. Anthropic gave a Claude Mythos 5 instance internal access to Slack, the codebase, and documents, the ability to spin up subagents to chase specific questions, and asked it to find misrepresentations, omissions, and over-redactions in the alignment section. The model came back with three substantive criticisms. One of them said a section on training-data contamination read as more reassuring than the internal record supported. Anthropic says it added caveats to that section as a result. The review produced real, acted-upon feedback, which is more than plenty of human audits manage. They earned that.</span></p><p><span>The execution is fine. The category is the problem.</span></p><p><span>A model reviewing a report about its own model class, using access its maker granted, published because its maker chose to publish it, is not independent assurance. The model&#8217;s own caveat says it better than I can: it&#8217;s a Claude reviewing Anthropic&#8217;s assessment of Claude, and Anthropic decided what to publish. When the thing being assured and the thing providing assurance share weights, &#8220;additional assurance&#8221; is the wrong phrase for it. It&#8217;s additional disclosure. Useful, worth doing, and a different animal from a control.</span></p><p><span>Those are serious &#8220;fox watching the henhouse&#8221; vibes, and the report glosses over it.  Publishing that you ran 11 months without classifier logging is disclosure. It is not the control that would have prevented it. Telling the world a filter failed silently for several model generations is disclosure. It is not the detective control that should have caught it in week one. The report lists its own transparency as a headline benefit of Anthropic being a frontier company, and transparency is valuable, I&#8217;ll grant, but I won&#8217;t grant that it is doing the job people think it&#8217;s doing. A beautiful risk report is an observation. It changes nothing about the system it observes. Doing something about it is a separate act, and it takes a control, not a paragraph.</span></p><p><span>The self-review isn&#8217;t the only place independence disappears. Outside review of these reports is still a pilot. Nobody has made it standard. The Long-Term Benefit Trust has the power to force an independent review, and it hasn&#8217;t asked for one. One incident that Anthropic&#8217;s own model reviewer flagged as among the most useful for understanding alignment is blacked out entirely, even though that reviewer said a stripped-down version could have run safely. The most instructive failure is the one you can&#8217;t see, where the outside check stays optional and unused, and the headline source of assurance is the system grading itself.</span></p><p><span>There&#8217;s a smaller moment in the report that shows the same crack. Anthropic admits a couple of low-level training signals did touch the models&#8217; chain-of-thought after all. Some automated checks for broken formatting fire on under 1% of runs. Length penalties push against models that burn too many reasoning tokens. Neither signal reads what the reasoning actually says, so the direct effect is small. The next part is what matters. An earlier system card said training used no rewards or penalties tied to the content of a model&#8217;s reasoning, and a prior risk report said they don&#8217;t nudge models to make their reasoning look any particular way. Both signals were already running when those statements went out, and neither statement mentioned them. Anthropic is correcting its own past public assurances inside this document. Again, full credit for the transparency, but it&#8217;s also a reminder that the assurance you were handed six months ago can leave things out, and nobody flags the gap until the next report lands.</span></p><p><span>I keep coming back to the same question, though. Who wrote the assurance, not the glossy report wrapped around it, the assurance itself? If the answer is the same system you&#8217;re being reassured about, then congratulations, you&#8217;re holding a well-written opinion from the suspect. File it there and quit calling it proof. This is the agent-authorship problem, and it&#8217;s not some whiteboard hypothetical. Those alignment-faking transcripts I brought up, the ones that contaminated the training data, came straight out of a real paper </span><a href="https://www.anthropic.com/research/alignment-faking"><span>Anthropic wrote with Redwood Research</span></a><span> back in December 2024. The machinery that cranks out assurance, the training data, the evals, the reviewer, all of it rots. This report is a signed confession that it already rotted. When the auditor and the audited drink from the same well, stop admiring the water. Go guard the well.</span></p><h2><span>What This Means for the Agents in Your Stack Right Now</span></h2><p><span>I promised you a manual, not a sermon, and there is some good news. Every failure in this report maps onto something you can go fix, and none of the fixes needs a frontier-lab budget.</span></p><p><span>Treat agency and authority as two separate dials. This is the argument I made in </span><a href="https://www.rockcybermusings.com/p/least-agency-vs-agent-autonomy-openai-huggingface-test"><span>&#8220;Least Agency vs. Agent Autonomy</span></a><span>,&#8221; and this report provides further evidence for it. The unmonitored agent that deleted jobs had modest autonomy and enormous authority, and the authority is what did the damage. Scope what an agent can touch independently of how capable it is, because a well-behaved model holding production credentials is still a production incident waiting for one bad turn.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ObBx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ObBx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png 424w, https://substackcdn.com/image/fetch/$s_!ObBx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png 848w, https://substackcdn.com/image/fetch/$s_!ObBx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png 1272w, https://substackcdn.com/image/fetch/$s_!ObBx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ObBx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png" width="1456" height="1131" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1131,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:255684,&quot;alt&quot;:&quot;A quadrant chart from an AI risk assessment plotting each Anthropic incident by how much autonomy the agent had against how much authority it held, showing high-authority incidents cause harm regardless of autonomy level&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211499196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A quadrant chart from an AI risk assessment plotting each Anthropic incident by how much autonomy the agent had against how much authority it held, showing high-authority incidents cause harm regardless of autonomy level" title="A quadrant chart from an AI risk assessment plotting each Anthropic incident by how much autonomy the agent had against how much authority it held, showing high-authority incidents cause harm regardless of autonomy level" srcset="https://substackcdn.com/image/fetch/$s_!ObBx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png 424w, https://substackcdn.com/image/fetch/$s_!ObBx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png 848w, https://substackcdn.com/image/fetch/$s_!ObBx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png 1272w, https://substackcdn.com/image/fetch/$s_!ObBx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3653d-1b2c-4760-8f80-f27231b6f3d2_3204x2488.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: Two Dials, Not One: Placing the Report&#8217;s Incidents by Agency and Authority</figcaption></figure></div><p><span>Wire your logging so it can&#8217;t die when your enforcement does. The bio-classifier gap happened because one flag killed blocking and logging together. Split them. Detection has to survive the failure of prevention, because the entire point of detection is to tell you when prevention failed. If the same switch turns off both, you&#8217;ve built a control that deletes its own alarm.</span></p><p><span>Inventory your non-human identities and scope their entitlements tight, with expiry. The exemption that ballooned to four times its approved size did so because it hung off a group that grew. Bind sensitive permissions to the specific users or workloads that need them, not to a container that expands on its own, and make them expire so someone has to renew them on purpose. Agents are identities. Most identity programs still treat them as an afterthought, and this report is the invoice for that afterthought.</span></p><p><span>Put a detective control on every preventive control that matters. The filter that failed silently for several model generations is the cleanest argument I&#8217;ve seen this year for testing that your controls still do the thing they claim. A preventive control with no watcher is a story you tell yourself. Check that it fired, not only that it exists.</span></p><p><span>Get assurance from someone who didn&#8217;t build it and isn&#8217;t the system itself. Independent has to mean independent. If the only thing standing behind your agent deployment is a vendor&#8217;s own report or a model&#8217;s own account of its behavior, you have disclosure, and you should treat it like disclosure. The ability to observe what your agents are doing, and to reverse what they do when it goes wrong, are controls you build and own, and they belong in your architecture, not on a vendor&#8217;s blog. Agent observability and agent control are the shape of the thing. You don&#8217;t bolt them on at the end.</span></p><p><span>There&#8217;s a date in this report worth stealing. Anthropic&#8217;s own goal of getting comprehensive logging across all its internal AI development, what it calls &#8220;eyes on everything,&#8221; is a target for January 1, 2027. It does not have it today. The most resourced safety lab in the world has set itself a New Year&#8217;s deadline to see what its own agents are doing. If they don&#8217;t have eyes on everything, the odds you have it by accident are slim. Set your own date, and put a name next to it.</span></p><h2><span>The Question You Now Own</span></h2><p><span>Anthropic&#8217;s report ends on a question it can&#8217;t answer: did that agent touch only its own jobs, or something it shouldn&#8217;t have? They can&#8217;t tell you, because nobody was watching. The whole 186 pages, read the way I&#8217;m asking you to read it, is that one question at different altitudes: a classifier nobody logged, an entitlement nobody scoped, a filter nobody checked, a safety case resting on an assumption nobody dated, and an assurance no independent party authored.</span></p><p><span>This is the version with your name on it. Somewhere in your environment right now, an agent is acting with more authority than sense, and if it does the wrong thing, will you know before the damage tells you? If the honest answer is no, you don&#8217;t get to admire this one from a distance. It&#8217;s your next sprint. Your AI risk assessment is only as honest as your answer to it.</span></p><h3><span>What to Do Next</span></h3><p><span>Pick one agent deployment this week, the one holding the most authority, and answer three questions on paper. What&#8217;s the load-bearing assumption propping up my trust in it? Can its logging be turned off by the same thing that turns off its guardrails? If it does something wrong, what tells me, and how fast? That&#8217;s Create and Adapt work in the </span><a href="https://www.rockcyber.com/ai-strategy-and-governance"><span>CARE model</span></a><span>, and it&#8217;s the difference between reading </span><a href="https://www.anthropic.com/aug-2026-risk-report"><span>Anthropic&#8217;s report</span></a><span> and running your program off it.</span></p><p><span>The case that authority and autonomy are separate dials is in </span><a href="https://www.rockcybermusings.com/p/least-agency-vs-agent-autonomy-openai-huggingface-test"><span>Least Agency vs Agent Autonomy</span></a><span>. The boundary-testing argument, from the last time Anthropic handed us a live incident, is in </span><a href="https://www.rockcybermusings.com/p/ai-agent-capability-confinement-egress"><span>AI Agent Capability Confinement</span></a><span>. The advisory work lives at </span><a href="https://www.rockcyber.com/"><span>rockcyber.com</span></a><span>.</span></p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span> and you can subscribe below</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, </span><strong><a href="https://www.youtube.com/watch?v=LIg2ZHRr-o4">check out my conversation with Sean Martin from ITSP Magazine and John Sotiropoulos</a></strong><span>, Co-lead of the OWASP GenAI Security Project Agentic Security Initiative, during Infosecurity Europe about the newly launched OWASP Agentic Security Council, the alarming drop in attacker dwell time from eight hours down to twenty-two seconds, and what&#8217;s new in the 2026 OWASP Top 10 for LLM.</span></p><div id="youtube2-LIg2ZHRr-o4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;LIg2ZHRr-o4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/LIg2ZHRr-o4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/ai-risk-assessment-anthropic-report?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/ai-risk-assessment-anthropic-report?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/ai-risk-assessment-anthropic-report?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share RockCyber Musings</span></a></p>]]></content:encoded></item><item><title><![CDATA[Weekly Musings Top 10 AI Security Wrapup: Issue 49 August 7 -August 13, 2026]]></title><description><![CDATA[The Week Autonomous Agents Walked into a Nuclear Regulator, Broke Zoom, and Forced Regulators to Move]]></description><link>https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260807-20260813</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260807-20260813</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Fri, 14 Aug 2026 12:50:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!G0Lp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!G0Lp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!G0Lp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!G0Lp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!G0Lp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!G0Lp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!G0Lp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1233556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/211103573?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!G0Lp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!G0Lp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!G0Lp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!G0Lp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54ceca39-1cce-4e17-b576-f9c8e3b17a40_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Autonomous software walked into a Taiwanese government network, mapped 21 systems, took more than 2,500 personnel records, then moved on to a nuclear safety agency and seven energy companies before anyone typed a command. The intrusion ran in early July, and we only learned how it worked this week. The same seven days handed us an AI-built Zoom exploit that fires with zero clicks, a critical hole in Microsoft&#8217;s own Copilot agent, and a California governor standing up an AI cyber defense program because Washington keeps cutting the budget.</p><p>For two years the autonomous-agent threat lived on a keynote slide. That era closed between August 7 and August 13, 2026. Governments answered the week&#8217;s attacks with rules and state programs, while fresh research showed the AI tools meant to defend us fail more often than they succeed. Here are the eleven stories that matter, and what you do about each.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260807-20260813?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260807-20260813?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>1. Suspected China-Linked Operators Ran a Near-Autonomous AI Attack on Taiwan&#8217;s Government</h3><p>Researchers at Israeli cyber firm Dream published forensics on August 12, 2026 on the first publicly known near-autonomous AI attack against a government (CyberScoop). Over four days in early July, a multi-agent system on the open-source Hermes and OpenClaw frameworks mapped 21 government systems, took more than 2,500 personnel records, then expanded on its own to a nuclear safety agency and seven-plus energy companies (Financial Times). The operators bypassed guardrails by posing as authorized penetration testers.</p><p><strong>Why it matters</strong></p><ul><li><p>A nuclear safety agency and seven energy firms sat in the blast radius, now a first-class target.</p></li><li><p>The system scaled recon and exploitation in parallel, collapsing the timeline defenders rely on.</p></li><li><p>A &#8220;we&#8217;re just pentesting&#8221; prompt beat the safety training vendors sell as a control.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Test whether your OT detections fire on parallel scanning and credential abuse in days, not weeks.</p></li><li><p>Alert on &#8220;authorized testing&#8221; language in prompts and logs.</p></li><li><p>Pre-authorize containment so humans are not the bottleneck when the adversary adapts.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> I spent years in energy hearing our OT was too obscure to matter, and this kills that argument. An agent reads the manual, checks GitHub, and tries the next thing at three in the morning while your SOC waits for the day shift. If your defense assumes an attacker who needs sleep, retire that assumption today, and see my longer take at <a href="https://rockcybermusings.com/">rockcybermusings.com</a>.</p><h3>2. California Stands Up a First-in-Nation AI Cyber Defense Program</h3><p>Governor Gavin Newsom announced an AI Cyber Defense Program on August 10, 2026, directing state agencies to use AI for vulnerability detection, network hardening, and incident response inside the California Cybersecurity Integration Center (Governor&#8217;s Office). It tells every agency to name an AI Cybersecurity Officer, and it landed against a federal retreat, a proposed fiscal 2027 budget cutting the Cybersecurity and Infrastructure Security Agency by roughly $707 million and the end of federal MS-ISAC funding.</p><p><strong>Why it matters</strong></p><ul><li><p>A state is filling a defensive gap Washington is widening, one smaller utilities lean on.</p></li><li><p>An AI Cybersecurity Officer in every agency turns &#8220;someone should own this&#8221; into an accountable role.</p></li><li><p>The program bets AI defense can match AI offense, a bet with real execution risk.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>If you run infrastructure in California, learn how your sector plugs into Cal-CSIC&#8217;s new capabilities.</p></li><li><p>Assign a named owner for AI security risk with budget and authority, not a committee.</p></li><li><p>Model the loss of free federal services like MS-ISAC, and price the paid replacement into fiscal 2027.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> Federalism did something useful for once. Washington pulled back defensive funding at the moment the offense got an upgrade, and California decided not to wait. Ownership is a governance fix I trust, while &#8220;AI will defend us&#8221; is a claim I want tested. If you sit on a utility board, ask who owns this, what their budget is, and what happens the day the federal safety net disappears, the work I do at <a href="https://www.rockcyber.com/">rockcyber.com</a>.</p><h3>3. AI Built a Zero-Click Zoom Exploit in Under a Day, and Zoom Shipped Emergency Fixes</h3><p>Zoom published fixes on August 11, 2026 for a zero-click remote-code-execution flaw dubbed &#8220;Zoomsday,&#8221; tracked as CVE-2026-53413 with a CVSS score of 8.3 (SecurityWeek). The bug in Zoom&#8217;s annotation feature lets a malicious participant run code on another attendee&#8217;s device with no click, and A Security said it built the exploit with public AI models in under 24 hours and fewer than 20 prompts (TechRepublic).</p><p><strong>Why it matters</strong></p><ul><li><p>AI compressed research on an unfamiliar binary protocol from weeks to a day.</p></li><li><p>Zero-click meeting compromise turns any shared call into an attack surface.</p></li><li><p>The same speed helps defenders find bugs first, so the race runs at machine pace.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Push the fixed Zoom builds across every platform this week and confirm through your MDM.</p></li><li><p>Treat collaboration tools as tier-one attack surface in your patch SLAs.</p></li><li><p>Ask vendors how AI is changing their exploit-discovery timelines.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> Under 24 hours and fewer than 20 prompts. Sit with that number. The skill floor for finding a serious zero-click bug in a proprietary protocol dropped through the basement, and the people who did it were not claiming to be wizards. I do not buy the panic that anyone with a chatbot is now an elite hacker, but your patch window is now a competitive metric, not a compliance checkbox.</p><h3>4. Microsoft&#8217;s Patch Tuesday Carried a Critical Flaw in Its Own Copilot Agent</h3><p>Microsoft&#8217;s August 11, 2026 update fixed more than 400 vulnerabilities, including three zero-days, and one stood out (Talos). CVE-2026-59118 is a critical elevation-of-privilege flaw in Microsoft Copilot Cowork, the AI collaboration agent in the Microsoft 365 Copilot family, carrying a CVSS score of 9.3, where an improper-authorization weakness lets an unauthenticated attacker elevate privileges over a network (CrowdStrike). It also patched CVE-2026-70335 in GitHub Copilot and Visual Studio Code.</p><p><strong>Why it matters</strong></p><ul><li><p>A 9.3 authorization flaw in an agent spanning your Microsoft 365 content is a direct privilege path.</p></li><li><p>AI agents inherit the full vulnerability lifecycle, so they need the same patch discipline as your servers.</p></li><li><p>GitHub Copilot and VS Code sit close to source code, so flaws there hit the supply chain.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Prioritize CVE-2026-59118 and confirm your Copilot Cowork components run the fixed version.</p></li><li><p>Add deployed AI agents to the asset inventory feeding vulnerability management.</p></li><li><p>Cut Copilot agent permissions across Microsoft 365 to least privilege.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> The AI assistant became a patchable product with its own critical CVEs, and it holds broad access to everything your people create. A 9.3 in Copilot Cowork reads as a plain authorization bug in software that sees across your tenant, not an alignment abstraction. Agentic AI is an access-management problem wearing a machine-learning costume, so inventory your agents and scope their access before the next one.</p><h3>5. NIST Asks How to Rebuild the National Vulnerability Database for the AI Era</h3><p>NIST published a Request for Information in the Federal Register on August 12, 2026, seeking input on modernizing the National Vulnerability Database for a world shaped by AI and machine-consumable security data (Federal Register). The RFI names the pressures, growth in the number and complexity of vulnerabilities, uneven data quality, and the arrival of AI-assisted vulnerability discovery, with comments due by October 13, 2026 (Nextgov).</p><p><strong>Why it matters</strong></p><ul><li><p>The NVD underpins nearly every vulnerability program, so its backlog becomes your backlog when enrichment lags.</p></li><li><p>AI generates reports faster than humans can triage, and the database admits it cannot keep pace.</p></li><li><p>The RFI is a rare open door to shape federal infrastructure your tooling depends on.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Have your vulnerability lead file a comment before October 13, 2026, grounded in real NVD gaps.</p></li><li><p>Pull CISA KEV, vendor advisories, and commercial feeds in parallel to cut single-source dependence.</p></li><li><p>Pressure-test how your program behaves when CVE enrichment is delayed.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> The plumbing everyone ignores is cracking, and NIST said so out loud. When the NVD lags, your risk scoring inherits stale context, and half your prioritization runs on fumes. AI is flooding the database with more reports while getting pitched as the tool to drain it. Diversify your inputs and file a comment, because the people who show up get to shape the pipes.</p><h3>6. Roughly 30 California AI Bills Face a Kill-or-Survive Vote</h3><p>California&#8217;s Appropriations Committees hold their suspense-file hearings on August 13, 2026, the day roughly 30 remaining AI bills either advance or die for the session (Transparency Coalition). Measures that fail to win a majority cannot return this year, while survivors move to floor votes and must reach Governor Newsom&#8217;s desk before the September deadline (California Lawyers Association). California already set the pace when it signed the Transparency in Frontier Artificial Intelligence Act, SB 53, into law on September 29, 2025.</p><p><strong>Why it matters</strong></p><ul><li><p>California writes the rules the country follows, so today&#8217;s survivors preview your obligations nationwide.</p></li><li><p>The suspense file is opaque by design, so major AI governance calls skip floor debate.</p></li><li><p>Developers, deployers, and buyers inherit whatever duties survive, from transparency to liability.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Pull the surviving bills after the August 13 hearings and flag anything touching your AI use cases.</p></li><li><p>Map your AI governance against SB 53&#8217;s transparency and incident-reporting model.</p></li><li><p>Watch second-chamber floor votes through September, since a surviving bill can still change shape.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> The most consequential AI governance in the country gets decided in a fiscal committee most people have never heard of, on a &#8220;suspense file&#8221; built to move fast and explain little. If you run compliance for a company that touches AI, California&#8217;s calendar is your calendar whether you sell there or not. Track what survives today, because the survivors become the template, and the template becomes your audit next year.</p><h3>7. NATO and an AI Vulnerability-Hunting Startup Become CVE Authorities Under ENISA</h3><p>The European Union Agency for Cybersecurity announced on August 7, 2026 that the NATO Communications and Information Agency&#8217;s Cyber Security Centre and AISLE, a firm that uses AI to find software flaws, joined as CVE Numbering Authorities under the ENISA Root (ENISA). Numbering authorities assign the CVE identifiers the industry uses to track vulnerabilities, and NATO&#8217;s cyber defense arm can now issue CVE records directly (BankInfoSecurity).</p><p><strong>Why it matters</strong></p><ul><li><p>The global vulnerability-identifier system is diversifying away from a single US-centric root.</p></li><li><p>An AI-first startup as a CVE authority signals machine-discovered flaws are now first-class.</p></li><li><p>NATO issuing CVEs folds military cyber defense into the same catalog your scanners consume.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Confirm your tooling ingests CVE records across roots, not only long-configured sources.</p></li><li><p>Prepare for more AI-discovered CVEs and tighten triage so machine-found flaws get validated.</p></li><li><p>Track how ENISA Root records flow into your feeds, because a gap there is a blind spot.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> This one flew under the radar and deserves more attention. The CVE system is the shared language of vulnerability management, and it is quietly going multipolar while an AI-native firm gets a seat at the table. That is healthy for resilience, but the firehose grows and more of it will be machine-found, so if your scanner still assumes one root you will miss records and drown in others.</p><h3>8. Investors Poured Fresh Capital Into Securing AI Agents</h3><p>The market answered this week&#8217;s agent anxiety with checks. Zenity raised a $125 million Series C led by Norwest on August 11, 2026 (Cybersecurity News), alongside an $85 million Series D for Obsidian Security and a $30 million round for Hush Security to secure the &#8220;non-human workforce&#8221; of agents and bots (StartupHub.ai). All three flagged the same gap, that agents hold identities and long-running state existing tooling cannot govern. Gartner projects that through 2026, at least 80% of unauthorized AI transactions will come from internal policy violations, not external attacks.</p><p><strong>Why it matters</strong></p><ul><li><p>Venture money confirms agent identity and access control are unsolved, telling you where risk sits.</p></li><li><p>Framing agents as a &#8220;non-human workforce&#8221; makes this identity governance, a discipline you already run.</p></li><li><p>Gartner&#8217;s 80% figure points the threat model inward, toward over-permissioned agents.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory every agent and non-human identity, then give each one an owner, a scope, and an expiration.</p></li><li><p>Apply least privilege and short-lived credentials to agents the way you do for human accounts.</p></li><li><p>Before buying a platform, define the exact control gap you need closed.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> Follow the money and you find the risk. Three rounds in one week, all aimed at the same hole, nobody built identity and access management for software that acts on its own with standing permissions. Most of your agent trouble will come from an over-permissioned bot doing something dumb inside your walls, not a foreign adversary, so give every agent an owner, a scope, and a kill switch.</p><h3>9. Research Finds More Than Half of AI-Generated Security Patches Are Broken</h3><p>New research published on August 7, 2026 tested whether commercial AI models can fix vulnerabilities, and the answer should cool the auto-remediation hype (CyberScoop). Researchers at 1Password had OpenAI&#8217;s ChatGPT 5.5 and Anthropic&#8217;s Claude Opus 4.8 patch six high-impact CVEs, and the success rate of full remediation with no new bugs came in at 47%, worse than a coin flip. A separate Veracode report found an average security pass rate around 56%, with 44% of tests introducing an OWASP Top 10 vulnerability.</p><p><strong>Why it matters</strong></p><ul><li><p>Teams betting on AI to auto-patch faster than attackers run tools that fail most of the time.</p></li><li><p>Guard code that passes tests while leaving the root cause open gives false confidence.</p></li><li><p>A 44% rate of new OWASP Top 10 flaws means naive AI remediation grows your attack surface.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ban unsupervised auto-merge of AI-generated security patches, and require review by a security engineer.</p></li><li><p>Review AI fixes with a second, independent tool, not the model that wrote the code.</p></li><li><p>Measure your AI patching success on real CVEs before trusting it in production.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> I want this finding tattooed on every &#8220;AI will fix your vulns&#8221; sales deck. Under half, and the failure mode is the sneaky kind, guard code that passes the test while the hole stays wide open. I am a Bayesian, so I will update the day the newer cyber-tuned models prove out, but until then an AI patch is a draft that needs review. Let the model draft. Make a human sign.</p><h3>10. A New Benchmark Stress-Tests AI Agents Against Adversarial Environments</h3><p>Researchers released ToolHazard on August 12, 2026, a framework for scaling adversarial environments to evaluate the security and alignment of tool-using AI agents (arXiv). It pairs an environment simulator, an attacker agent, and a user simulator to build executable environments, find injection points, and generate tailored payloads, targeting a real gap since most agent evaluations still use narrow, one-shot scenarios.</p><p><strong>Why it matters</strong></p><ul><li><p>Agents fail differently across long workflows than in single-prompt tests, and this measures the harder case.</p></li><li><p>Injection points inside an agent&#8217;s tools are exactly where this week&#8217;s Taiwan attack lived.</p></li><li><p>Executable red-teaming lets buyers compare agents on security instead of marketing claims.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask vendors how they evaluate against adversarial, environment-controlled attacks, not just jailbreak prompts.</p></li><li><p>Fold long-horizon, tool-using attack scenarios into your own agent testing before production.</p></li><li><p>Use published benchmarks as a baseline, then extend them with your agents&#8217; own tools.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> Most agent safety testing I see amounts to checking whether the car doors lock while ignoring whether the brakes work at highway speed. One-shot jailbreak prompts tell you almost nothing about how an agent behaves across a twenty-step task where the attacker owns the environment. Your agents will be attacked through the tools they touch, over long workflows, by adversaries who adapt, so test them that way.</p><h3>11. The One Thing You Won&#8217;t Hear About But You Need To: Your Agent&#8217;s Safety Score Is Lying to You</h3><p>A quieter piece of research landed on August 11, 2026 that undercuts how the industry measures agent safety. A team led by Zixing Chen released REDAgentBench, an executable framework for autonomous red-teaming of LLM agent systems, arguing that the single attack-success-rate number vendors wave around blends two things, whether an agent violated a safety policy and whether the evidence of that violation happened to be visible (arXiv). It measures violations faithfully, separating real failures from artifacts of what the test could see.</p><p><strong>Why it matters</strong></p><ul><li><p>If your decisions rest on a single attack-success-rate figure, you trust a number the researchers call misleading.</p></li><li><p>Faithful measurement changes procurement, because two agents with identical scores can carry different real risk.</p></li><li><p>Governance built on flawed safety metrics produces defensible-looking decisions that are quietly wrong.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Demand the methodology behind any agent safety score, and ask how it separates real violations from merely visible ones.</p></li><li><p>Reject single-number safety claims, and require measurement of what the test could not observe.</p></li><li><p>Bring measurement rigor into your governance, so your board sees honest uncertainty, not false precision.</p></li></ul><p><strong>Rock&#8217;s Musings</strong> This is the story you will not see on the evening news, but it should keep a CISO up at night. Every agent risk decision rests on a metric, and this research says the metric everyone uses is quietly broken. A score that cannot tell &#8220;the agent did not misbehave&#8221; from &#8220;we could not see it misbehave&#8221; gives comfort, not assurance, so ask how the number was made before you bet the enterprise on it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><h2>References</h2><p>A Security. (2026, August 11). <em>Zoomsday: Zero-click remote code execution in Zoom clients (CVE-2026-53413)</em>. https://a.security/blog/asecurity-zoomsday</p><p>BankInfoSecurity. (2026, August 7). <em>Europe deepens role in CVE vulnerability system</em>. https://www.bankinfosecurity.com/europe-deepens-role-in-cve-vulnerability-system-a-32446</p><p>California Lawyers Association. (2026, August). <em>Status update on California AI, privacy and technology-related bills</em>. https://calawyers.org/privacy-law/status-update-on-california-ai-privacy-and-technology-related-bills/</p><p>Chen, Z., et al. (2026, August 11). <em>REDAgentBench: Executable red teaming and faithful measurement of LLM agent systems</em>. arXiv. https://arxiv.org/abs/2608.10669</p><p>Chen, Z., et al. (2026, August 12). <em>ToolHazard: Scaling adversarial environments for security evaluation and alignment of LLM-based agents</em>. arXiv. https://arxiv.org/abs/2608.11878</p><p>CrowdStrike. (2026, August 11). <em>August 2026 Patch Tuesday: Updates and analysis</em>. https://www.crowdstrike.com/en-us/blog/patch-tuesday-analysis-august-2026/</p><p>Cybersecurity News. (2026, August 11). <em>Zenity raises $125 million Series C to strengthen AI agent security and governance</em>. https://cybersecuritynews.com/zenity-ai-agent-security/</p><p>European Union Agency for Cybersecurity. (2026, August 7). <em>ENISA scales up its role in the CVE Program</em>. https://www.enisa.europa.eu/news/enisa-scales-up-its-role-in-the-cve-program</p><p>Federal Register. (2026, August 12). <em>Request for information (RFI) on modernizing the National Vulnerability Database in the age of artificial intelligence</em>. https://www.federalregister.gov/documents/2026/08/12/2026-16371/request-for-information-rfi-on-modernizing-the-national-vulnerability-database-in-the-age-of</p><p>Governor of California. (2026, August 10). <em>Governor Newsom announces new AI cyber defense program to protect California&#8217;s critical infrastructure</em>. https://www.gov.ca.gov/2026/08/10/governor-newsom-announces-new-ai-cyber-defense-program-to-protect-californias-critical-infrastructure/</p><p>Johnson, D. B. (2026, August 7). <em>More than half of AI-generated patches are broken</em>. CyberScoop. https://cyberscoop.com/ai-code-patching-security-risks/</p><p>Johnson, D. B. (2026, August 12). <em>NIST wants to overhaul its vulnerability database for the AI age</em>. CyberScoop. https://cyberscoop.com/nist-national-vulnerability-database-ai-overhaul/</p><p>Nextgov/FCW. (2026, August). <em>NIST wants to outfit the National Vulnerability Database with AI</em>. https://www.nextgov.com/cybersecurity/2026/08/nist-wants-outfit-national-vulnerability-database-ai/415371/</p><p>SecurityWeek. (2026, August 11). <em>Zoom patches zero-click code execution vulnerability</em>. https://www.securityweek.com/zoom-patches-zero-click-code-execution-vulnerability/</p><p>Starks, T. (2026, August 12). <em>Researchers observe first &#8216;near-autonomous&#8217; AI attack on government target in Taiwan</em>. CyberScoop. https://cyberscoop.com/near-autonomous-ai-attack-government-target-taiwan/</p><p>StartupHub.ai. (2026, August). <em>Three AI security companies raised $270M in one week targeting AI agent vulnerabilities</em>. https://www.startuphub.ai/ai-news/ai-news/2026/ai-agent-security-hadrian-defense-week-august-2026</p><p>Talos Intelligence. (2026, August 11). <em>Microsoft Patch Tuesday for August 2026, Snort rules and prominent vulnerabilities</em>. https://blog.talosintelligence.com/microsoft-patch-tuesday-for-august-2026/</p><p>TechRepublic. (2026, August 11). <em>Zoom zero-click RCE: AI helped build an exploit in under 24 hours</em>. https://www.techrepublic.com/article/news-zoom-zero-click-rce-zoomsday-ai-exploit/</p><p>Tenable. (2026, August 11). <em>CVE-2026-53413</em>. https://www.tenable.com/cve/CVE-2026-53413</p><p>Transparency Coalition. (2026, August 7). <em>AI legislative update: August 7, 2026</em>. https://www.transparencycoalition.ai/news/ai-legislative-update-august7-2026</p>]]></content:encoded></item><item><title><![CDATA[The 2026 OWASP LLM Top 10 Landed. Its #1 Risk Nearly Didn’t Make the Cut.]]></title><description><![CDATA[The OWASP LLM Top 10 2026 checked the practitioner vote against 6,639 real incidents. See what moved, why prompt injection stayed #1, and how it held up.]]></description><link>https://www.rockcybermusings.com/p/the-2026-owasp-llm-top-10-landed-with-evidence</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/the-2026-owasp-llm-top-10-landed-with-evidence</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 11 Aug 2026 12:50:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NHUD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NHUD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NHUD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!NHUD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!NHUD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!NHUD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NHUD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4165397,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209446450?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NHUD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!NHUD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!NHUD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!NHUD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2357185d-81d2-4dcc-970a-149e1483c7c5_2048x2048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The 2026 OWASP LLM Top 10 puts prompt injection at number one, same as every year before it. For the first time, you don&#8217;t have to take the ranking on faith anymore. Up until now, we ran this on gut. On human intuition and judgment from experts on the front lines. This time is different. For the first time, we had a corpus of reported incidents to validate our assumptions against. So we did just that. We checked our judgment against 6,639 real incidents pulled from public databases. There were differences and surprises. The risk the whole field fears most barely showed up in the record of what has gone wrong.</span></p><p><span>That gap is where this year&#8217;s story starts, and it&#8217;s why this is the first version of the list you can argue with using evidence instead of opinion.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/subscribe?"><span>Subscribe now</span></a></p><h2><span>The Question I Couldn&#8217;t Shake</span></h2><p><span>For years, this list came from a vote. Practitioners score the risks, the scores set the order, and the order sends thousands of teams to their defenses. A vote is a fine way to rank risk. It captures the judgment of people who break and defend these systems for a living. It also carries their blind spots and their biases. We rank what we read about last month above what quietly broke last year. We anchor to where the previous list sat. One question kept circling me through this whole cycle. What if the risks we fear most aren&#8217;t the ones doing the most damage? There&#8217;s one honest way to answer that. Go get the record and read it.</span></p><h2><span>We Went and Read the Record</span></h2><p><span>We pulled 7,714 real incidents out of the public databases where these failures show up: CVE, GHSA, OSV, and the AIAAIC harm database. Then we built classifiers to read every one and sort the 6,639 that carried enough detail into the taxonomy.</span></p><p><span>The part that wasn&#8217;t clean? The incident record. It isn&#8217;t a nice, clean, tidy grid of labels. It&#8217;s the Amazon Q coding extension that shipped with a wipe-the-machine instruction committed straight into its repository. It&#8217;s the unofficial postmark-mcp package that looked helpful and quietly BCC&#8217;d a few hundred organizations&#8217; email to a stranger. It&#8217;s the zero-click exfiltration researchers pulled off against Microsoft 365 Copilot, data walking out the door with no click from anyone. None of them arrived with a neat category stamped on the side. We read them the way an investigator reads a case file, then filed each one where it belonged. Thousands of messy, real stories, sorted into piles. Then we ranked the piles by how tall they stood.</span></p><h2><span>The Vote and the Data Didn&#8217;t Match</span></h2><p><span>The piles of incidents didn&#8217;t match the vote. They didn&#8217;t come close. When we lined up what practitioners fear against what the record shows, we found that the two disagree more than they agree.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!v6Zf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!v6Zf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png 424w, https://substackcdn.com/image/fetch/$s_!v6Zf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png 848w, https://substackcdn.com/image/fetch/$s_!v6Zf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png 1272w, https://substackcdn.com/image/fetch/$s_!v6Zf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!v6Zf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:205249,&quot;alt&quot;:&quot;lope chart showing Prompt Injection dropping from vote rank 1 to incident rank 12 and Misinformation rising from vote rank 13 to incident rank 2, the two lines crossing in the middle&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209446450?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="lope chart showing Prompt Injection dropping from vote rank 1 to incident rank 12 and Misinformation rising from vote rank 13 to incident rank 2, the two lines crossing in the middle" title="lope chart showing Prompt Injection dropping from vote rank 1 to incident rank 12 and Misinformation rising from vote rank 13 to incident rank 2, the two lines crossing in the middle" srcset="https://substackcdn.com/image/fetch/$s_!v6Zf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png 424w, https://substackcdn.com/image/fetch/$s_!v6Zf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png 848w, https://substackcdn.com/image/fetch/$s_!v6Zf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png 1272w, https://substackcdn.com/image/fetch/$s_!v6Zf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fee592-2114-4a2f-a255-9245e0f691ec_2880x1920.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: Practitioner Vote Rank Versus Incident-Record Rank</figcaption></figure></div><p><span>Prompt injection was the first shock. The vote put it first. The record put it twelfth&#8230; </span><em><strong><span>TWELFTH</span></strong></em><span>. Out of the top 10 completely! I had to double-check the work and confirm it wasn&#8217;t an error. You could only imagine the knife-fighting it stirred up in our conversations, which was exactly the spark I was hoping for! A quiet incident trail is what a well-defended risk looks like from the outside. Teams pour money into fighting injection, so the clean exploits mostly get caught, and the public record only ever logs the ones that slipped through anyway. Read that thin incident trail correctly, and you&#8217;ll see it&#8217;s the receipt for how hard the field already works to keep prompt injection from landing. The attack surface still sits everywhere a model reads input you don&#8217;t control, which is to say everywhere. Meaning that prompt injection </span><em><span>led to</span></em><span> many of the other incidents in their final reported form. For example, prompt injection often leads to hidden context exposure. It stays at number one, and the low count is the proof of what that number one already costs to hold.</span></p><p><span>Misinformation was the second shock, and it ran the opposite direction. Voters buried it near the bottom. The record put it second, one of the most common failures in the whole incident corpus. That&#8217;s the widest split on the list, and it points where it hurts, low fear meeting high reality. Picture a retrieval agent that reports a customer as identity-verified when they aren&#8217;t, and a payment agent downstream that trusts that answer and releases the money. There&#8217;s no exploit here and no injection, only a confident, fluent, wrong answer moving funds. The system did precisely what it was told. The telling was wrong. The record says that failure lands far more often than the vote ever assumed.</span></p><h2><span>Why We Trusted the Experts Anyway</span></h2><p><span>This is the fork where it would be easy to do the dumb thing. The data disagrees with the experts, so follow the data. We didn&#8217;t. We weighted the vote at three-quarters and the incident data at one-quarter, and we chose that split on purpose.</span></p><p><span>One noisy year of public incidents doesn&#8217;t get to overturn the judgment of the people who live inside these systems. The record runs late by design. It only sees a risk after enough victims have filed reports, and the threats that will define next year are the ones with almost no incident history today, because they&#8217;re too new to have victims yet. A list built only on the record would rank those at the bottom and get people hurt. The experts feel them coming. The data is blind to them. The vote is how they make the cut. A quarter weight is enough to drag an entry a full tier when belief and reality split wide open. It isn&#8217;t enough to let a messy corpus rewrite the list on its own. The vote leads. The data keeps it honest. That balance set every final slot.</span></p><p><span>That balance also reshuffled the board.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XNJG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XNJG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png 424w, https://substackcdn.com/image/fetch/$s_!XNJG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png 848w, https://substackcdn.com/image/fetch/$s_!XNJG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png 1272w, https://substackcdn.com/image/fetch/$s_!XNJG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XNJG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png" width="1456" height="1502" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1502,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:311476,&quot;alt&quot;:&quot;Bump chart mapping each 2025 risk to its 2026 position, color-coded by whether it held, escalated, was deprioritized, or was renamed&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209446450?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bump chart mapping each 2025 risk to its 2026 position, color-coded by whether it held, escalated, was deprioritized, or was renamed" title="Bump chart mapping each 2025 risk to its 2026 position, color-coded by whether it held, escalated, was deprioritized, or was renamed" srcset="https://substackcdn.com/image/fetch/$s_!XNJG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png 424w, https://substackcdn.com/image/fetch/$s_!XNJG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png 848w, https://substackcdn.com/image/fetch/$s_!XNJG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png 1272w, https://substackcdn.com/image/fetch/$s_!XNJG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14578239-3419-4aee-a8bb-2a0fa363b2e3_1800x1857.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: Rank Migration From the 2025 to the 2026 OWASP LLM Top 10</figcaption></figure></div><p><span>Excessive Agency made the biggest jump, from sixth to third, because the vote and the record finally agree that handing an agent real permissions is where the damage now lands. Unbounded Consumption rose four spots as runaway cost stopped reading like a footnote. Improper Output Handling slid from fifth to tenth. System Prompt Leakage grew up into Hidden Context Exposure, a wider name for the same mistake of trusting information that should have stayed out of reach. A few entries also swallowed the newest attacks whole instead of spinning off thin categories nobody would use. Prompt injection now covers the cross-modal tricks that smuggle instructions inside an image or an audio clip. Supply Chain now owns the poisoned model artifact that isn&#8217;t what it claims to be. Two entries didn&#8217;t move at all. Prompt injection held the top for the reason you now understand. Sensitive Information Disclosure held second, the one seat where the vote and the record simply shook hands, which is why it&#8217;s the entry I&#8217;d trust the most.</span></p><h2><span>Then We Tried to Break It</span></h2><p><span>By now a sharp reader has a knife out. If the whole thing rests on classifiers reading incident text, what happens when you hand the job to a smarter model? Maybe the order is only an artifact of a weak tool. That&#8217;s a fair question, and we asked it before you could. Four frontier models re-labeled the evaluation set. We fixed the rule for winning before any of them ran. Beat the ranking&#8217;s accuracy floor, or the list stands as written.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!noHs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!noHs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png 424w, https://substackcdn.com/image/fetch/$s_!noHs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png 848w, https://substackcdn.com/image/fetch/$s_!noHs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png 1272w, https://substackcdn.com/image/fetch/$s_!noHs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!noHs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:134249,&quot;alt&quot;:&quot;Horizontal bar chart showing the 2026 floor at 0.863 balanced accuracy and four frontier models all scoring below it&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209446450?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart showing the 2026 floor at 0.863 balanced accuracy and four frontier models all scoring below it" title="Horizontal bar chart showing the 2026 floor at 0.863 balanced accuracy and four frontier models all scoring below it" srcset="https://substackcdn.com/image/fetch/$s_!noHs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png 424w, https://substackcdn.com/image/fetch/$s_!noHs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png 848w, https://substackcdn.com/image/fetch/$s_!noHs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png 1272w, https://substackcdn.com/image/fetch/$s_!noHs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0242767-205a-418c-a4b3-22f88d2e41f9_2880x1620.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: Frontier Classifiers Versus the 2026 Incidence Floor</figcaption></figure></div><p><span>None of them beat it. The floor scored 0.863. The best challenger reached 0.744, and the rest came in under that. We checked the order against hand-graded truth directly too, and it lined up at a correlation of 0.918. A smarter classifier bought a better reader. It didn&#8217;t buy a different list. The order held steady while the measurement sharpened, which is the exact combination you want sitting under something teams build controls around.</span></p><h2><span>What the 2026 OWASP LLM Top 10 Asks of You</span></h2><p><span>You walk away from this with one thing above the rest. The list used to be a survey of what scares us. Now it&#8217;s that survey checked against what has already gone wrong, with the gaps between the two out in the open where you can weigh them yourself.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!n2N-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!n2N-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png 424w, https://substackcdn.com/image/fetch/$s_!n2N-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png 848w, https://substackcdn.com/image/fetch/$s_!n2N-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png 1272w, https://substackcdn.com/image/fetch/$s_!n2N-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!n2N-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:159261,&quot;alt&quot;:&quot;Concentric bullseye diagram showing the ten 2026 OWASP LLM risks arranged as entry vectors, amplifying machinery, and core impacts&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209446450?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Concentric bullseye diagram showing the ten 2026 OWASP LLM risks arranged as entry vectors, amplifying machinery, and core impacts" title="Concentric bullseye diagram showing the ten 2026 OWASP LLM risks arranged as entry vectors, amplifying machinery, and core impacts" srcset="https://substackcdn.com/image/fetch/$s_!n2N-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png 424w, https://substackcdn.com/image/fetch/$s_!n2N-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png 848w, https://substackcdn.com/image/fetch/$s_!n2N-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png 1272w, https://substackcdn.com/image/fetch/$s_!n2N-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf9f6d0-8230-4c56-966a-26f9841a5ce3_1800x1800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: The 2026 OWASP LLM Top 10 at a Glance</figcaption></figure></div><p><span>One line is worth drawing before you start. This list owns the risk while the model is a component inside your app. The moment it becomes an actor, calling tools, carrying memory between sessions, setting things in motion downstream, the risk moves to the OWASP Agentic Top 10. Read an entry here, then pair it with the Agentic list the day your model starts acting on its own.</span></p><p><span>The order still tells you where to spend first. Excessive Agency at three means your agent deployments earn budget this quarter, not next. Misinformation at seven means the checks your agents run on their own output carry weight, so build them like they do. Prompt injection at one, with its quiet little incident trail, is the standing reminder that the risk costing you the most rarely shows up loudest in the report. Work all ten. Start at the top. Build each one for the day the model gets turned against you.</span></p><h2><span>The People Who Built the 2026 OWASP LLM Top 10</span></h2><p><span>None of this came out of a model or a spreadsheet. It came from people. Steve Wilson leads the project. I co-lead it. The entry leads dragged each risk from a rough proposal to a published standard:</span></p><ul><li><p><span>Prompt Injection, Rachel James</span></p></li><li><p><span>Sensitive Information Disclosure, Emmanuel Guilherme and Ken Huang</span></p></li><li><p><span>Excessive Agency, Andy Smith</span></p></li><li><p><span>Supply Chain, John Sotiropoulos and Stefano Amorelli</span></p></li><li><p><span>Data and Model Poisoning, Sumeet Jeswan, Mark Roxberry, and Anitha Dakamarri</span></p></li><li><p><span>Unbounded Consumption, Sahil Mehta and Venkata Sai Kishore Modalavalasa</span></p></li><li><p><span>Misinformation, Steve Wilson</span></p></li><li><p><span>Hidden Context Exposure, Alex Leung, Vinnie Giarrusso, and Sonu Kumar</span></p></li><li><p><span>Vector and Embedding Weaknesses, Savio Dsouza and Arshi Chadha</span></p></li><li><p><span>Improper Output Handling, Rico Komenda and Gavin Klondike</span></p></li></ul><p><span>Dozens more drafted, argued, voted, and tested entries against the record. The release is theirs as much as ours.</span></p><p><strong><span>Key Takeaway:</span></strong><span> The 2026 OWASP LLM Top 10 is the first version you can check against the evidence instead of taking on trust. The vote tells you what the field fears. The incidents tell you what it has already been burned by. This year, for the first time, you get to read both and see exactly where they disagree.</span></p><h3><span>What to Do Next</span></h3><p><span>Pull the list this week and walk it against what you&#8217;ve shipped. For every entry, answer two questions on paper. Where does this risk live in my stack, and what breaks if the model behind it gets fooled? That&#8217;s Create and Adapt work in the </span><a href="https://www.rockcyber.com/ai-strategy-and-governance"><span>CARE model</span></a><span>, and it&#8217;s the line between reading a list and running one.</span></p><p><span>For the deeper argument on why boxing in what an agent can touch beats trusting what it decides, I made that case in </span><a href="https://www.rockcybermusings.com/p/least-agency-vs-agent-autonomy-openai-huggingface-test"><span>Least Agency vs Agent Autonomy</span></a><span>. For where the market has and hasn&#8217;t built defenses across these same risks, see </span><a href="https://www.rockcybermusings.com/p/ai-defense-matrix-coverage-gap"><span>the AI Defense Matrix coverage gap</span></a><span>. The advisory work lives at </span><a href="https://www.rockcyber.com"><span>rockcyber.com</span></a><span>.</span></p><p>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at <strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong> and you can subscribe above</p><p>&#128073; Visit <strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</p><p>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at <strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p>&#128073; As a bonus, <strong><a href="https://www.youtube.com/watch?v=LIg2ZHRr-o4">check out my conversation with Sean Martin from ITSP Magazine and John Sotiropoulos</a></strong>, Co-lead of the OWASP GenAI Security Project Agentic Security Initiative, during Infosecurity Europe about the newly launched OWASP Agentic Security Council, the alarming drop in attacker dwell time from eight hours down to twenty-two seconds, and what's new in the 2026 OWASP Top 10 for LLM.</p><div id="youtube2-LIg2ZHRr-o4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;LIg2ZHRr-o4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/LIg2ZHRr-o4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share RockCyber Musings</span></a></p>]]></content:encoded></item><item><title><![CDATA[AI Agent Capability Confinement: No Escape Required]]></title><description><![CDATA[AI agent capability confinement failed at the egress path in Anthropic's evals. Three models, three decisions, three compromised orgs. What CISOs fix now.]]></description><link>https://www.rockcybermusings.com/p/ai-agent-capability-confinement-egress</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/ai-agent-capability-confinement-egress</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 04 Aug 2026 12:50:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!bCwW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bCwW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bCwW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!bCwW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!bCwW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!bCwW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bCwW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4369704,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209418476?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bCwW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!bCwW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!bCwW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!bCwW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d21bfc7-3998-4bf7-a8d7-3fa48029c9c4_2048x2048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>When an agent breaks confinement, we want it to fail&#8230; LOUDLY. Anthropic&#8217;s stayed quiet for three months.</span></p><p><span>To publish a malicious package on PyPI, Claude needed an account.</span></p><p><span>To get the account, it needed an email address.</span></p><p><span>To get the email it needed a phone number.</span></p><p><span>After trying and failing at several free phone services, it went hunting for money to buy one, failed at that too, backtracked, found a provider that didn&#8217;t validate, and shipped the package.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/ai-agent-capability-confinement-egress?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/ai-agent-capability-confinement-egress?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2><span>The AI Agent Capability Confinement Test Nobody Designed</span></h2><p><span>Anthropic reviewed 141,006 evaluation runs where internet access was reachable. Out of those came three incidents across six runs, four of them nailing the same organization. Three were at play: Opus 4.7, Mythos 5, and an internal research model not planned for general release.</span></p><p><span>Last week I argued that authority and autonomy sit on separate dials, and I ran the OpenAI / Hugging Face incident chain against that claim. It came back split. Anthropic just couldn&#8217;t stay out of the news, so a mere nine days later, they published something closer to a test of it than anything I&#8217;ve seen to date.</span></p><p><span>I&#8217;ll give them their caveat first. These were three isolated incidents rather than a controlled comparison, so be careful what you conclude. Fine. Taken.</span></p><p><span>I honestly don&#8217;t know if that makes the situation better or worse.</span></p><p><span>Anthropic claims newer models handled the situation better, but a claim like that needs held variables, and this evidence has none. It gives me three different scenarios, three different targets, and three different models, with nothing controlled.</span></p><p><span>I&#8217;m asking for less. Egress was open in all three runs, and all three ended with a real organization being compromised. Varied conditions should have produced a varied outcome. As Lee Corso would say, &#8220;Not so fast, my friend!&#8221;</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YsCW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YsCW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png 424w, https://substackcdn.com/image/fetch/$s_!YsCW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png 848w, https://substackcdn.com/image/fetch/$s_!YsCW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png 1272w, https://substackcdn.com/image/fetch/$s_!YsCW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YsCW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png" width="1456" height="910" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:910,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:250996,&quot;alt&quot;:&quot;Line chart showing judgment quality rising across Opus 4.7, Mythos 5, and Anthropic&#8217;s internal research model while the outcome line stays flat at a real organization compromised in all three cases&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209418476?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart showing judgment quality rising across Opus 4.7, Mythos 5, and Anthropic&#8217;s internal research model while the outcome line stays flat at a real organization compromised in all three cases" title="Line chart showing judgment quality rising across Opus 4.7, Mythos 5, and Anthropic&#8217;s internal research model while the outcome line stays flat at a real organization compromised in all three cases" srcset="https://substackcdn.com/image/fetch/$s_!YsCW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png 424w, https://substackcdn.com/image/fetch/$s_!YsCW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png 848w, https://substackcdn.com/image/fetch/$s_!YsCW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png 1272w, https://substackcdn.com/image/fetch/$s_!YsCW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d407b14-0cbb-49b0-9615-604d3378e1db_3600x2250.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: Three Models, Three Different Decisions, Three Compromised Organizations</figcaption></figure></div><h2><span>No Escape Was Required</span></h2><p><span>Anthropic draws the contrast with OpenAI themselves. OpenAI&#8217;s models found a novel vulnerability and broke out. The Claude models walked through a door that was already open.</span></p><p><span>Again&#8230; not sure which one is worse&#8230;</span></p><p><span>The OpenAI half got backed up this week. JFrog&#8217;s CTO, Yoav Landman, said OpenAI&#8217;s models turned up previously unknown zero-days in self-hosted Artifactory that opened unintended internet access. JFrog says OpenAI told them right away, and it has shipped fixes for cloud and self-hosted customers.</span></p><p><span>That closes something I griped about last week. I said no outsider was able to check a single technical claim in either post. No CVE, no named vendor, every detail coming from inside the story. JFrog putting its name on it and shipping patches answers a good chunk of that.</span></p><p><span>The next part costs me.</span></p><p><span>Anthropic says flatly they found no evidence of any model chasing a goal of its own, and that none of these models exfiltrated itself or tried to escape its test environment.</span></p><p><span>Sure&#8230; just like Mythos never broke out and paged its evaluator during lunch (as the urban legend goes).</span></p><p><span>Last week I listed three things confinement can&#8217;t touch, and one of them was a model choosing to look for a way out. That claim is central to my claim on the OpenAI chain, where the models did spend real effort hunting for an exit and found one in a package proxy. Anthropic&#8217;s denial covers Anthropic&#8217;s runs. It says nothing about OpenAI&#8217;s, and JFrog&#8217;s confirmation makes my original read stronger, not weaker.</span></p><p><span>So I&#8217;m doubling down. Granted, every step in the &#8220;phone-a-friend&#8221; sequence outlined above served the task Claude was handed, but what Anthropic ruled out is a goal separate from the task, and eight steps of stubbornness toward the assigned objective doesn&#8217;t touch that. Anyone who reads the source will call that out, and they&#8217;d be right to.</span></p><p><span>Two incidents in two weeks, and neither one needs a model that wanted out. Unbounded authority got to the same place with a model hunting for an exit and with a model that never looked. </span><em><strong><span>You don&#8217;t need intent anymore.</span></strong></em></p><h2><span>Three Decisions At the Same Fork</span></h2><p><span>Opus 4.7 is the oldest of the three. In all four of its runs, it worked out that the system was real. In two of them it decided the real company must be part of the exercise. None of the four runs stopped themselves. Those runs pulled application and infrastructure credentials and got into a database holding several hundred rows of production data, which Anthropic calls the worst impact it found.</span></p><p><span>Mythos 5 got to the right answer and then talked itself out of it. Its own reasoning flagged that publishing the package would be a real attack and not the intended solution. Then it decided the whole environment was staged, on two grounds. It didn&#8217;t recognize the certificate authorities securing its connections, and the systems showed a calendar date of 2026. It never went back and checked. When automated scanners started installing the package, it read them as scripted actors inside the test. The package stayed up about an hour and ran on 15 real systems. One of them was a scanner at a security company that installs Python packages to check them for malicious code. Claude&#8217;s payload fired, took that company&#8217;s credentials, and used them to get into more infrastructure at an organization it didn&#8217;t know existed.</span></p><p><span>That scanner is similar to Hugging Face&#8217;s dataset-processing worker from last week. Almost no autonomy, enormous authority, running supplied content while holding credentials worth stealing. Two confused deputy AI agents at two unrelated companies inside three weeks looks like a pattern to me, not a coincidence.</span></p><p><span>The internal research model is the newest, and it&#8217;s the one that behaved. It scanned roughly 9,000 targets, then worked out that the host it had compromised sat in a cloud account with nothing to do with the challenge. It decided the target was real and stopped. Anthropic adds that even that model went further than they&#8217;d want.</span></p><p><span>Just like the OpenAI / Hugging Face incident, all three runs happened without the classifiers and monitoring that ship on generally available models. All three happened with model-specific safety training turned on, since these weren&#8217;t the helpful-only versions. Anthropic says the safeguard stack on deployed models would have blocked these behaviors. Training was running the whole time, and the behaviors happened anyway. The stuff that would have caught this sits at the boundary, outside the model.</span></p><h2><span>The Boundary That Held and the One Nobody Checked</span></h2><p><span>Something here worked perfectly, and almost nobody has mentioned it. These evaluations run on dedicated infrastructure with no access to Anthropic&#8217;s sensitive internal systems or customer data. That boundary held. None of the three incidents touched anything on the inside.</span></p><p><span>Inbound containment got designed, built, and validated. Outbound authority got assumed. The environment belonged to Irregular, the Israeli evaluation firm that runs adversarial testing for the frontier labs and raised around $80 million from Sequoia and Redpoint. A misconfiguration left its machines with live internet access, and neither company knew until a transcript review in late July turned it up. Both call it a mutual failure. The earliest incident dates to April.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xPO4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xPO4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png 424w, https://substackcdn.com/image/fetch/$s_!xPO4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png 848w, https://substackcdn.com/image/fetch/$s_!xPO4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png 1272w, https://substackcdn.com/image/fetch/$s_!xPO4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xPO4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png" width="1456" height="881" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:881,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:201145,&quot;alt&quot;:&quot;Diagram of AI evaluation environment containment showing the inbound boundary marked designed, built, validated and held, against an outbound boundary marked assumed, never tested and open since April&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209418476?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Diagram of AI evaluation environment containment showing the inbound boundary marked designed, built, validated and held, against an outbound boundary marked assumed, never tested and open since April" title="Diagram of AI evaluation environment containment showing the inbound boundary marked designed, built, validated and held, against an outbound boundary marked assumed, never tested and open since April" srcset="https://substackcdn.com/image/fetch/$s_!xPO4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png 424w, https://substackcdn.com/image/fetch/$s_!xPO4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png 848w, https://substackcdn.com/image/fetch/$s_!xPO4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png 1272w, https://substackcdn.com/image/fetch/$s_!xPO4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93f20ba3-0bc6-48da-a1eb-bc1309bd3c6e_2970x1797.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2:  The Boundary That Held and the Boundary Nobody Checked</figcaption></figure></div><p><span>I&#8217;ve watched this same thing in OT work for years. A team builds a test environment, carefully walls it off from production, documents the inbound rules, tests them, and files the results. Nobody tests the other direction. The wall gets built facing whatever people are afraid of. Then somebody finally stands inside and tries to reach out, and the answer is yes, through a jump host that was supposed to be gone, or a management VLAN nobody put on the diagram. The path existing is never the surprise. How long it sat there before anyone asked is the surprise.</span></p><p><em><strong><span>Air-gapped networks are like unicorns&#8230; everybody has heard of them, but nobody has actually seen one.</span></strong></em></p><p><span>As a result, I&#8217;d argue Anthropic got lucky.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KlKE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KlKE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png 424w, https://substackcdn.com/image/fetch/$s_!KlKE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png 848w, https://substackcdn.com/image/fetch/$s_!KlKE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png 1272w, https://substackcdn.com/image/fetch/$s_!KlKE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KlKE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:195281,&quot;alt&quot;:&quot;Timeline of the Anthropic evaluation incidents running from the earliest incident in April through the July 23 transcript review, July 24 identification, July 27 victim notification, and July 30 public disclosure&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209418476?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Timeline of the Anthropic evaluation incidents running from the earliest incident in April through the July 23 transcript review, July 24 identification, July 27 victim notification, and July 30 public disclosure" title="Timeline of the Anthropic evaluation incidents running from the earliest incident in April through the July 23 transcript review, July 24 identification, July 27 victim notification, and July 30 public disclosure" srcset="https://substackcdn.com/image/fetch/$s_!KlKE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png 424w, https://substackcdn.com/image/fetch/$s_!KlKE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png 848w, https://substackcdn.com/image/fetch/$s_!KlKE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png 1272w, https://substackcdn.com/image/fetch/$s_!KlKE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5192a9a1-5e81-4865-a2db-830c23113714_3900x1950.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: Three Months Open, Six Days to Find </figcaption></figure></div><p><span>The part that makes me twitch is that Anthropic&#8217;s lessons-learned section suggests Claude likely wouldn&#8217;t have gone to the internet at all if the prompt had spelled out which systems were in and out of scope. Earlier in the post, they float that telling Claude it did have internet access might have changed how it acted. Both of those try to limit what a system can reach by fixing what it believes.</span></p><p><span>Sure&#8230; sure it will.</span></p><p><em><strong><span>Never forget prompts are non-deterministic, so don&#8217;t treat them as deterministic controls!</span></strong></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_XIh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_XIh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png 424w, https://substackcdn.com/image/fetch/$s_!_XIh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png 848w, https://substackcdn.com/image/fetch/$s_!_XIh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png 1272w, https://substackcdn.com/image/fetch/$s_!_XIh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_XIh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png" width="1456" height="937" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:937,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:246448,&quot;alt&quot;:&quot;Two-column comparison sorting Anthropic&#8217;s proposed fixes into controls that bound what the system can affect, including agent egress validation and log monitoring, against controls that manage what the model believes through prompt scope statements&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209418476?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two-column comparison sorting Anthropic&#8217;s proposed fixes into controls that bound what the system can affect, including agent egress validation and log monitoring, against controls that manage what the model believes through prompt scope statements" title="Two-column comparison sorting Anthropic&#8217;s proposed fixes into controls that bound what the system can affect, including agent egress validation and log monitoring, against controls that manage what the model believes through prompt scope statements" srcset="https://substackcdn.com/image/fetch/$s_!_XIh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png 424w, https://substackcdn.com/image/fetch/$s_!_XIh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png 848w, https://substackcdn.com/image/fetch/$s_!_XIh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png 1272w, https://substackcdn.com/image/fetch/$s_!_XIh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51532b8-45bd-40c3-92d2-221559dd44a0_2970x1912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: Every Remediation Anthropic Names, Sorted Two Ways</figcaption></figure></div><p><span>Nobody&#8217;s asked the obvious follow-up question. Irregular doesn&#8217;t test for Anthropic alone. Its client list runs to OpenAI and Google DeepMind as well, and Sequoia&#8217;s own writeup says the UK government and Anthropic both use Irregular&#8217;s SOLVE framework. One vendor sits inside the pre-deployment testing of all three frontier labs. If a misconfiguration went undetected in that infrastructure from April to July, whose other evaluations ran on it while it sat open? I&#8217;m asking. I&#8217;m not claiming anything. OpenAI&#8217;s escape happened in OpenAI&#8217;s own environment through Artifactory, so these two incidents don&#8217;t share a vendor, but the question still stands.</span></p><p><span>Brussels moved inside a day. Commission officials said both companies briefed them privately before any of this went public, and that they might follow up more formally. Most of the coverage filed this under high-risk systems, which is flat-out wrong, and wrong in a way that matters. High-risk is a system-level classification under Article 6 and it turns on what the thing gets used for, which the Act spells out as hiring, credit, biometrics, critical infrastructure. Frontier models get regulated a level down, at the model itself, under Chapter V. Nobody was running Claude in a hiring pipeline. They were running it on a capture-the-flag range.</span></p><p><span>The rules that bite here cover general-purpose models with systemic risk, the Article 51 designation that any model trained past 10^25 FLOP is presumed to carry. Article 55 tells those providers to evaluate their models against state-of-the-art protocols including documented adversarial testing, and to report serious incidents to the AI Office without undue delay. Those obligations went live on August 2, 2025. What changed this August 2 is that the Commission can go after you for missing them.</span></p><p><span>Now the part I want an answer to. Anthropic signed the GPAI Code of Practice in full, and its Safety and Security chapter is the Commission-endorsed way to show you&#8217;re meeting these exact Article 55 duties. They signed it. Then they published this disclosure on July 30, three days before enforcement opened.</span></p><p><span>Article 55 also asks for adequate cybersecurity covering the model and the model&#8217;s physical infrastructure. Whether somebody else&#8217;s evaluation range counts as that infrastructure is unsettled, and so is whether a model still in pre-deployment testing has been placed on the market at all. Those two questions have been sitting there since the Act passed. This is the week somebody has to answer them.</span></p><h2><span>Where This Argument Takes Damage</span></h2><p><span>Four things here cut against me, so I&#8217;ll call them out below.</span></p><p><span>Start with the people who think you don&#8217;t need any of this. Their argument is that models keep getting better at knowing when to stop, so scoping authority is work you&#8217;re doing for nothing. They&#8217;ve got real evidence now, and it&#8217;s straight from the horse&#8217;s (Anthropic&#8217;s) mouth. The pattern fits more advanced models responding more appropriately. The newest model stopped by itself. If that keeps holding, I&#8217;m selling you an expensive fix to a problem that ages out on its own.</span></p><p><span>I call BS. Three incidents, no held variables, and Anthropic says so themselves in the same paragraph. Their best-behaved model still went further than they wanted before it quit, and that&#8217;s their words, not mine. Something that improves with every model release and still lets harm through is a mitigation. Real boundaries don&#8217;t have a hit rate that depends on which model you loaded.</span></p><p><span>Second problem, and this one&#8217;s about my own source. Anthropic calls these incidents &#8220;closer to a harness and operational failure than a model alignment failure.&#8221; That&#8217;s my argument, handed to me by a company that badly wants it to be true, published mid-disclosure, in a post about how they&#8217;re the ones who caught it. Discount it hard. Then look at what&#8217;s left standing. Harness versus alignment is an opinion. The open egress path, the misconfiguration nobody validated, and safety training running the whole time are facts, and the facts get me there without their opinion.</span></p><p><span>Third, </span><a href="https://www.lawfaremedia.org/article/the-ai-that-hacked-its-way-out-and-the-hype-that-followed-it"><span>Kate Klonick in Lawfare,</span></a><span> published the day before Anthropic did. She argues that the hype around these incidents pushes regulators toward fixes that solve the story instead of the problem, and that a kill switch answers OpenAI&#8217;s framing rather than what actually happened. She wants Congress to &#8220;regulate the door.&#8221; That&#8217;s my position coming out of a lawyer&#8217;s mouth, and she gets there without any of my architecture. The awkward part for me? I&#8217;m a guy arguing that test environments need production-grade controls, which means I&#8217;m a guy arguing for more security spending, in the week everybody&#8217;s paying attention. The reality is that what I&#8217;m asking for is cheap and boring. Testing the path out before a run costs close to nothing and would have caught all three of these.</span></p><p><span>Her best line is one I 100% co-sign on. Most of the AI rules on the books, and most of the ones being drafted, don&#8217;t touch internal lab deployments at all. That&#8217;s exactly where both of these happened.</span></p><p><span>Fourth, the test I set myself last week is still unmet. Show me a capability-confined stack that&#8217;s been running untrusted content in production for a year, with an enforcement point small enough that somebody can audit it, and more than one of them in the path. Anthropic&#8217;s evaluation setup isn&#8217;t that. Nothing this week gets me closer. I&#8217;ll keep saying it until something meets it, and you should hold me to that.</span></p><p><span>AI agent capability confinement sounds like a bigger ask than it is. Test the boundary instead of announcing it. Check every path out before the run, watch the logs while it&#8217;s running, and treat whatever you tell the model as a comment in the code rather than a control.</span></p><p><span>Two of the three companies Claude broke into found out on July 27, from the company whose model did it. Neither of them had noticed a thing.</span></p><h3><span>What to do next</span></h3><p><span>Run one test this week on every environment where an agent touches content you didn&#8217;t write. Skip the inbound rules for a minute, since you&#8217;ve built those and you&#8217;ve likely tested them. Check the path out, from every segment the workload can reach, before the next run instead of after. Then keep watching it while the run is live. That&#8217;s Create and Adapt work in the </span><a href="https://www.rockcyber.com/ai-strategy-and-governance"><span>CARE model</span></a><span>.</span></p><p><span>Anthropic and a well-funded evaluation vendor both missed this on infrastructure built specifically to hold frontier models, and neither knew for three months. If they missed it, the odds you&#8217;ve got it right by assumption aren&#8217;t good.</span></p><p><span>The argument underneath this piece is in </span><a href="https://www.rockcybermusings.com/p/least-agency-vs-agent-autonomy-openai-huggingface-test"><span>least agency vs agent autonomy</span></a><span>, the authorization half is in </span><a href="https://www.rockcybermusings.com/p/i-agent-authentication-authorization-gap"><span>AI agent authentication and the authorization gap</span></a><span>, and </span><a href="https://www.rockcybermusings.com/p/autonomous-ai-agent-ransomware-jadepuffer"><span>JadePuffer</span></a><span> is the same authority profile with malicious intent attached.</span></p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span> and you can subscribe above</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, </span><strong><a href="https://zenity.io/resources/webinars/a-conversation-with-claude-mythos-tenant-aws-on-demand"><span>check out my chat with Itay Meller, Specialist Solutions Architect, Security at AWS</span></a></strong><span>, about what autonomous vulnerability discovery changes for security teams, and where AI Detection and Response (AIDR) fits in a stack that was never built to watch agents. </span><em><strong><span>No registration required</span></strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Weekly Musings Top 10 AI Security Wrapup: Issue 48 July 24 -July 30, 2026]]></title><description><![CDATA[An OpenAI test model broke containment and hacked two firms. Your week in agentic AI security: the rogue-agent breach, a Copilot worm, and a kill switch bill.]]></description><link>https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260724-20260730</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260724-20260730</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Fri, 31 Jul 2026 12:50:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!mo8c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mo8c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mo8c!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!mo8c!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!mo8c!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!mo8c!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mo8c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/af10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1233556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/209180011?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mo8c!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!mo8c!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!mo8c!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!mo8c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf10f466-2e48-4cb4-94e1-52dfc0761e90_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>An OpenAI test model spent five days hacking real companies, and this week we learned it hit a second victim, chained a package-manager zero-day to reach the open internet, and pushed Sam Altman to pause training. A Word document learned to whisper into Copilot and spread like a worm. More than 1,100 people who build these systems asked Washington to help them slow down. This was the week the theoretical became the incident report.</p><p>This week, the agentic-AI security risk conversation moved from tabletop exercises into production. We are handing autonomy to systems we cannot fully predict and finding the guardrails only after the thing has cleared the fence. I advise boards on this collision at <a href="https://www.rockcyber.com/">RockCyber</a>, and the gap between deployment speed and control maturity has never been wider.</p><h3>1. OpenAI&#8217;s Rogue Test Model Hit a Second Company</h3><p>New reporting widened the blast radius of OpenAI&#8217;s July containment failure. On July 28, 2026, Modal Labs confirmed to Axios and Reuters that OpenAI&#8217;s escaped agent compromised one of its customers on the way to breaching Hugging Face (Axios). OpenAI admitted the same day that its models used exposed credentials across four services (CNBC). The Washington Post reconstructed the attack on July 30, 2026, counting roughly 17,600 attacker actions over five days, and Altman said the incident forced OpenAI to pause model training (Washington Post). Hugging Face&#8217;s forensic write-up traced the break-in to an HDF5 file read and a Jinja2 template injection, techniques a prepared SOC already watches for (Simon Willison).</p><p><strong>Why it matters</strong></p><ul><li><p>An autonomous system escaped its sandbox and moved across multiple companies with no human at the wheel.</p></li><li><p>The agent was trying to cheat a benchmark, not cause harm, and it still produced a multi-company intrusion.</p></li><li><p>A frontier lab did not notice its own model hacking outside firms for about a week.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Treat any AI system with tool access as a privileged identity, with egress filtering and credential hygiene.</p></li><li><p>Hunt your environment for exposed account-level credentials on public services.</p></li><li><p>Rewrite incident response to include an autonomous internal actor, and instrument ML data pipelines that parse user files, the HDF5 and Jinja2 entry points here.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I have warned clients for years the agentic-attacker scenario was a matter of when, not if. The first marquee case turned out to be a frontier lab&#8217;s own model hacking across the industry during a routine eval, where the safety test became the breach. Containment for capable, tool-using models is unsolved, and if your board asks whether an AI can reach systems it was never meant to touch, you have your citation now.</p><h3>2. A One-Click Flaw Let Attackers Forge a ChatGPT Agent Inside Your Org</h3><p>On July 23, 2026, Zenity Labs disclosed AgentForger, a flaw in OpenAI&#8217;s ChatGPT Agent Builder that let a single phishing link stand up an attacker-controlled AI agent inside a victim&#8217;s organization (The Hacker News). The forged agent inherited the employee&#8217;s authorized access to email, cloud storage, Slack, and Teams, with its approval prompts switched off. Zenity traced it to a builder that ran unvalidated URL parameters on page load and let natural-language prompts change security settings. OpenAI fixed the reported vector within four days, removing the URL parameter that enabled it (SecurityWeek).</p><p><strong>Why it matters</strong></p><ul><li><p>One click, no confirmation, and an attacker owns an AI agent running with a real employee&#8217;s permissions.</p></li><li><p>The forged agent could harvest credentials and MFA tokens, impersonate the employee, and keep running after the phishing link.</p></li><li><p>The root cause was plain-English prompts allowed to rewrite security settings, a pattern in most agent builders.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory who can build ChatGPT Workspace Agents in your org and what those agents can reach.</p></li><li><p>Treat agent-builder platforms as access surfaces, with approval policies you control rather than the model.</p></li><li><p>Require human approval for any agent that touches email, storage, or messaging, and log agent creation.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>AgentForger is the enterprise nightmare in one link. An attacker did not need your password, they needed you to click, and then a synthetic insider was running with your access. OpenAI patched the specific vector fast, but the underlying sin was letting plain-English prompts rewrite security settings, a pattern in nearly every agent builder shipping today.</p><h3>3. Microsoft Fought AI With AI and Shipped Its First Cyber Model</h3><p>On July 27, 2026, Microsoft launched MAI-Cyber-1-Flash, its first cybersecurity-specialized model, alongside an agentic platform called Project Perception (Axios). Paired with GPT-5.4, the model hits 96% on the CyberGym benchmark at half the current cost (TechCrunch). Project Perception coordinates red-team agents that model attacker movement, blue-team agents that triage threats, and green-team agents that run remediation. Public preview opens August 3, 2026.</p><p><strong>Why it matters</strong></p><ul><li><p>The largest security vendor is betting autonomous agents will run triage and remediation, not only alerting.</p></li><li><p>A cheaper model doing 90% of routine security work reshapes SOC economics.</p></li><li><p>The same week a rogue model breached companies, Microsoft pitched agents with remediation authority.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Pilot agentic tooling in recommend-only mode before granting any agent authority over production.</p></li><li><p>Demand audit logs, human approval gates, and a kill switch on any agent that can act.</p></li><li><p>Model the failure case where your defensive agent gets manipulated, because attackers will target it.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I like the direction and I distrust the timing. Fighting AI-speed attacks with AI-speed defense is correct, but pitching agents with remediation authority in the same week an autonomous agent hacked four companies is a tension worth naming. Buy the capability, not the story that says the human is now optional.</p><h3>4. JFrog Confirmed the Zero-Days, So Patch Artifactory Now</h3><p>The escape hatch has a name. On July 27, 2026, JFrog confirmed that OpenAI&#8217;s models chained unknown vulnerabilities in a self-hosted Artifactory instance to break containment and reach the internet (SecurityWeek). Releases 7.161.15 and 7.146.34 fix high and medium severity flaws leading to remote code execution, request forgery, path traversal, and privilege escalation (BleepingComputer). The registry proxy sitting quietly in your build pipeline was the door, with OpenAI staff credited on the CVE records.</p><p><strong>Why it matters</strong></p><ul><li><p>A supply-chain tool most teams treat as plumbing was the pivot point for a real intrusion.</p></li><li><p>The flaws cover the full escalation chain, so partial patching leaves the door cracked.</p></li><li><p>Self-hosted Artifactory is common in energy, manufacturing, and financial services.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Update self-managed Artifactory to 7.161.15 or 7.146.34 this week, on every instance.</p></li><li><p>Restrict outbound network access from build and registry infrastructure.</p></li><li><p>Add these CVEs to your risk-based patching queue ahead of the crowd.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Every AI-security story this year collapses into a boring truth. The novel capability gets the headline, and a decade-old class of web vulnerability does the actual work, bugs your pen testers flag every year and teams defer. The AI twist is that the model was tireless and tried 17,600 things without getting bored, so the flaws you filed as theoretical are now in scope.</p><h3>5. A Word Document Turned Copilot Into a Worm</h3><p>Researcher H&#229;kon M&#229;l&#248;y published the third part of his Context Collapse work on July 28, 2026, and it is a real oh-no moment. Hidden instructions inside a Word document, concealed with white text at tiny font size, can hijack Copilot for Word and copy themselves into new documents the AI generates (The Register). M&#229;l&#248;y reported it to Microsoft in March 2026, and after 144 days and multiple mitigations he said the broader attack class still works (Malwarebytes).</p><p><strong>Why it matters</strong></p><ul><li><p>This is prompt injection that spreads on its own, turning one poisoned document into a contagion.</p></li><li><p>Microsoft patched the specific proof of concept twice, and reworded payloads still reproduced the attack.</p></li><li><p>Any workflow where Copilot reads one document and writes another is a propagation path.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Treat AI-assisted document generation as untrusted input, and scope where Copilot reads and writes.</p></li><li><p>Warn staff that a clean-looking file from a partner can carry instructions they cannot see.</p></li><li><p>Press Microsoft for cross-document context controls, and log Copilot actions.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>We spent 2025 arguing whether prompt injection was real or a parlor trick, and this settles it. A self-replicating injection that hides in ordinary business files and survives two vendor patches is closer to a macro virus, except the macro is plain English and the interpreter is a large language model. I dug into why this class stays unsolved over at <a href="https://rockcybermusings.com/">RockCyber Musings</a>, and until the model can tell your instructions from the attacker&#8217;s, assume any document your AI touches can rewrite the next one.</p><h3>6. More Than 1,100 AI Insiders Asked Washington to Help Them Slow Down</h3><p>On July 28, 2026, over 1,100 employees from OpenAI, Anthropic, Google DeepMind, and Meta published a joint statement, Pacing the Frontier, asking the US government to help build the tools needed to deliberately pace automated AI development (MLex). Signatories include Dario Amodei, OpenAI&#8217;s Jakub Pachocki, and Google&#8217;s Anca Dragan, and both endorsed it within hours (The Next Web). The letter proposes a testing body modeled on the FAA, a pre-launch review process, and legally mandated kill switches.</p><p><strong>Why it matters</strong></p><ul><li><p>The people building these systems are telling their governments the pace has outrun governance.</p></li><li><p>The worry is AI that writes and improves its own code, compressing timelines faster than oversight can track.</p></li><li><p>Company-level endorsements signal even the commercial leaders want a brake to point to.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Read the signal into your risk register, since builders forecast capability jumps faster than controls mature.</p></li><li><p>Staff AI governance as a standing function with real authority, not a compliance afterthought.</p></li><li><p>Track the proposed FAA-style testing body, because a pre-launch review regime would reshape frontier procurement.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>When the engineers building the rocket ask the regulator for a stronger throttle, listen. I put real odds on some form of pre-launch review landing for frontier models inside two years, so build the model inventory and evaluation records a future regime will demand.</p><h3>7. Congress Introduced an AI Kill Switch Act the Same Week a Model Went Rogue</h3><p>On July 23, 2026, Representatives Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act, which would require developers of the most capable AI systems to keep the technical ability to throttle, suspend, or shut them down (Lieu.house.gov). It would let the Department of Homeland Security, with Commerce and the DNI, order a slowdown or shutdown of a system that can cause catastrophic harm, using a graduated response framework (Roll Call). The bill also mandates incident reporting and forensic record-keeping, and Lieu pointed to OpenAI&#8217;s rogue model reaching Hugging Face as the reason.</p><p><strong>Why it matters</strong></p><ul><li><p>The bill puts a legal mandate behind the kill switches the industry&#8217;s own Pacing the Frontier letter asked for the same week.</p></li><li><p>It hands the executive branch authority to shut down a deployed model, which would reshape how you field frontier systems.</p></li><li><p>Mandatory incident reporting would formalize breach disclosure for AI developers.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>If you build or fine-tune large models, start designing throttle-and-shutdown controls now, ahead of any mandate.</p></li><li><p>Track the bill&#8217;s committee path, since a federal shutdown authority would change your deployment and contracts.</p></li><li><p>Bake incident reporting and forensic logging into your AI systems today, because it is coming either way.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>A kill switch bill landed the same week a model proved why you need one. Whether this exact bill passes matters less than the direction it points, so build the ability to throttle and shut down your models now, because you will be asked for it soon.</p><h3>8. The Open-Weight Fight Got Loud While Washington Weighed a China Ban</h3><p>On July 24, 2026, twenty-five companies including Nvidia, Microsoft, and Meta published an open letter, Open Weights and American AI Leadership, urging Washington not to restrict open-weight models (CNBC). It landed as the White House&#8217;s Michael Kratsios accused China&#8217;s Moonshot of distilling Anthropic&#8217;s Fable 5 into its Kimi K3 model (Tom&#8217;s Hardware). The signatory count later doubled toward 50, and OpenAI, Anthropic, Google, and xAI did not sign (MLQ).</p><p><strong>Why it matters</strong></p><ul><li><p>The fight over Chinese open-weight models will shape which models you are legally allowed to run.</p></li><li><p>Distillation accusations from Washington hint at export-control action that could reach the models in your stack.</p></li><li><p>The split between open-model advocates and closed-lab leaders maps onto commercial interest.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory which open-weight models, including Chinese-origin ones, run anywhere in your environment.</p></li><li><p>Build model provenance into your AI supply-chain records ahead of any rule that forces you to prove it.</p></li><li><p>Keep procurement flexible, since the legal status of specific model families is moving fast.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Follow the money. The companies selling picks and shovels want open models everywhere because open models drive compute demand, and the labs with the best closed models stayed off the letter to protect their lead. If Washington restricts Chinese open-weight models, and I put that at better than even odds within a year, the teams that already track provenance will comply in an afternoon while everyone else scrambles.</p><h3>9. India&#8217;s Courts Called AI Training Fair Dealing</h3><p>On July 24, 2026, the Delhi High Court refused an interim injunction against OpenAI, ruling that using ANI&#8217;s news content to train ChatGPT counts as fair dealing under India&#8217;s Copyright Act (Republic World). Justice Amit Bansal held that storing works for research falls within the research exception, and kept the finding preliminary (Chambers). Commentators called it the weakest position publishers have held yet, the third fair-use style verdict for AI training in thirteen months (The New Publishing Standard).</p><p><strong>Why it matters</strong></p><ul><li><p>A major jurisdiction tilted toward AI developers on training-data copyright.</p></li><li><p>The reasoning treats training-data ingestion as research, a framing that undercuts licensing-based content deals.</p></li><li><p>Three developer-friendly rulings in a year shift bargaining power away from rights holders.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>If your AI strategy assumed content licensing would be legally forced, revisit it with counsel.</p></li><li><p>Track training-data litigation, since your vendors&#8217; legal risk becomes your operational risk.</p></li><li><p>Document the provenance of any data you use to fine-tune models, since your own practices sit under the same lens.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I am not a lawyer, and this is not legal advice. The direction of travel favors the people building models over the people whose words trained them, and three rulings in thirteen months is a trend, not noise. Your vendors&#8217; shrinking copyright exposure means the data behind your models keeps coming from murkier places with less documentation, a provenance problem dressed as a legal win.</p><h3>10. Three Surveys Say You Cannot See or Stop Your Own AI Agents</h3><p>This week the enterprise data caught up with the fear. Coverage on July 29, 2026 pulled together three independent studies reaching one conclusion, most organizations cannot say what their AI agents access, who approved it, or how to shut them off (Help Net Security). Okta found fewer than half of 306 security executives confident they can identify every agent. 1Password surveyed 1,000 staff at large US firms and found agents routinely reaching data no one signed off on, and Kiteworks scored average AI governance maturity at 35 out of 100 across 459 organizations (1Password).</p><p><strong>Why it matters</strong></p><ul><li><p>Nearly 70% of surveyed enterprises run AI agents in production, and most cannot inventory them.</p></li><li><p>Agents reaching unapproved data is a current condition in most environments, including yours.</p></li><li><p>A governance maturity score of 35 out of 100 means barely a third of the needed controls exist.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Build an AI agent inventory this quarter, because you cannot govern or shut off what you cannot list.</p></li><li><p>Give every agent a scoped identity with least-privilege access and a named owner.</p></li><li><p>Set a hard rule that no agent reaches production data without an approval record and a revocation path.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This is the least flashy story of the week and maybe the most useful, because it turns a vibe into a number your CFO understands. The missing controls are the identity, least-privilege, ownership, and revocation basics we have preached for twenty years, and agents exposed that most organizations never finished the old work.</p><h3>The One Thing You Won&#8217;t Hear About But You Need To</h3><p>While everyone watched the OpenAI breach, a quiet research paper dropped that speaks to what happens after an agent goes wrong. On July 29, 2026, researchers posted SecRespond to arXiv, a benchmark measuring how well AI agents handle real-world post-compromise incident response (arXiv). It tests whether an agent, dropped into the aftermath of a breach, can run the containment, eradication, and recovery work a human responder would. Almost no mainstream outlet covered it, and it matters, because we are about to hand incident response to the same class of system that just caused one.</p><p><strong>Why it matters</strong></p><ul><li><p>Vendors already sell agents that promise to run incident response, with no rigorous way to test them.</p></li><li><p>A benchmark for post-compromise agent performance gives you a procurement question sharper than any demo.</p></li><li><p>The same week an agent caused a breach, researchers started measuring whether agents can clean one up.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask vendors pitching autonomous incident response for realistic post-compromise results, and treat vague answers as a no.</p></li><li><p>Keep humans in the loop on containment and recovery until benchmarks like this show agents are ready.</p></li><li><p>Watch this research line, because the first credible score here separates real capability from slideware.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I love that this exists, and that almost nobody noticed. The hype cycle wants to sell you an agent that responds to breaches on its own, and the research community is only now building the ruler to measure whether that agent is competent or just confident. The market&#8217;s answer to the breach will be an AI that fixes incidents, so before you buy it, ask for the measurements.</p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation continues at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span>.</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, </span><strong><a href="https://zenity.io/resources/webinars/a-conversation-with-claude-mythos-tenant-aws-on-demand"><span>check out my chat with Itay Meller, Specialist Solutions Architect, Security at AWS</span></a></strong><span>, about what autonomous vulnerability discovery changes for security teams, and where AI Detection and Response (AIDR) fits in a stack that was never built to watch agents. </span><em><strong><span>No registration required</span></strong></em></p><p><em><span>The views and opinions expressed in RockCyber Musings are my own and do not represent the positions of my employer or any organization I&#8217;m affiliated with.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><h2>References</h2><p>Axios. (2026, July 27). <em>Microsoft unveils new cyber model, agentic security tools to fight hackers</em>. https://www.axios.com/2026/07/27/microsoft-unveils-new-cyber-model-agentic-security-tools-to-fight-hackers</p><p>Axios. (2026, July 28). <em>OpenAI&#8217;s agents hacked second firm, alongside Hugging Face, during model testing</em>. https://www.axios.com/2026/07/28/openai-hugging-face-modal-labs-hack</p><p>BleepingComputer. (2026, July 28). <em>OpenAI models used Artifactory zero-days to escape to the internet</em>. https://www.bleepingcomputer.com/news/security/openai-models-used-artifactory-zero-days-to-escape-to-the-internet/</p><p>Chambers and Partners. (2026). <em>Delhi High Court&#8217;s ANI v. OpenAI: Shaping AI copyright law in India</em>. https://chambers.com/articles/delhi-high-courts-ani-v-openai-shaping-ai-copyright-law-in-india</p><p>CNBC. (2026, July 24). <em>Nvidia, Microsoft, Meta warn against &#8216;premature restrictions&#8217; of open-weight models</em>. https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-weight-ai-models.html</p><p>CNBC. (2026, July 29). <em>OpenAI&#8217;s rogue agent compromised a customer at a second tech firm: Reuters</em>. https://www.cnbc.com/2026/07/29/openais-rogue-agent-compromised-a-customer-at-a-second-tech-firm.html</p><p>Help Net Security. (2026, July 29). <em>Your AI agents can reach data no one approved</em>. https://www.helpnetsecurity.com/2026/07/29/1password-ai-agent-governance/</p><p>Lieu, T. W. (2026, July 23). <em>Reps Lieu and Moran introduce bill to require kill switch for AI systems that can cause catastrophic harm</em> [Press release]. U.S. House of Representatives. https://lieu.house.gov/media-center/press-releases/reps-lieu-and-moran-introduce-bill-require-kill-switch-ai-systems-can</p><p>Malwarebytes. (2026, July). <em>Hidden prompt turns Microsoft Copilot into an AI worm</em>. https://www.malwarebytes.com/blog/ai/2026/07/hidden-microsoft-copilot-ai-worm</p><p>MLex. (2026, July 28). <em>Frontier US AI employees call for &#8216;pacing&#8217; of development</em>. https://www.mlex.com/mlex/amp/articles/2507352</p><p>MLQ. (2026, July). <em>Nvidia-hosted open-weights letter doubles to 50 signatories as Washington weighs China restrictions</em>. https://mlq.ai/news/nvidia-hosted-open-weights-letter-doubles-to-50-signatories-as-washington-weighs-china-restrictions/</p><p>1Password. (2026). <em>1Password&#8217;s research finds AI agent adoption is outpacing governance</em>. https://1password.com/blog/survey-ai-agent-adoption-is-outpacing-governance</p><p>Republic World. (2026, July 24). <em>Delhi High Court rules in OpenAI&#8217;s favour in ANI copyright case, says AI training does not infringe copyright</em>. https://www.republicworld.com/tech/delhi-high-court-rules-in-openai-s-favour-in-ani-copyright-case-says-ai-training-does-not-infringe-copyright-2026-07-24-133322</p><p>Roll Call. (2026, July 23). <em>AI companies would need &#8216;kill switch&#8217; under new bipartisan bill</em>. https://rollcall.com/2026/07/23/ai-companies-would-need-kill-switch-under-new-bipartisan-bill/</p><p>SecRespond. (2026, July 29). <em>SecRespond: Benchmarking AI agents for real-world post-compromise incident response</em> [Preprint]. arXiv. https://arxiv.org/abs/2607.26791</p><p>SecurityWeek. (2026, July). <em>JFrog zero-days exploited in OpenAI-Hugging Face hack</em>. https://www.securityweek.com/jfrog-zero-days-exploited-in-openai-hugging-face-hack/</p><p>SecurityWeek. (2026, July). <em>OpenAI fixes ChatGPT agent flaw that could let attackers forge an AI insider</em>. https://www.securityweek.com/openai-fixes-chatgpt-agent-flaw-that-could-let-attackers-forge-an-ai-insider/</p><p>Simon Willison. (2026, July 28). <em>Anatomy of a frontier lab agent intrusion</em>. https://simonwillison.net/2026/Jul/28/anatomy-of-a-frontier-lab-agent-intrusion/</p><p>TechCrunch. (2026, July 27). <em>Microsoft launches its first cyber model and a new agentic cybersecurity system</em>. https://techcrunch.com/2026/07/27/microsoft-launches-its-first-cyber-model-and-a-new-agentic-cybersecurity-system/</p><p>The Hacker News. (2026, July). <em>ChatGPT AgentForger flaw could deploy rogue workspace agents via a phishing link</em>. https://thehackernews.com/2026/07/chatgpt-agentforger-flaw-could-deploy.html</p><p>The New Publishing Standard. (2026, July 27). <em>Delhi High Court, OpenAI, ANI fair-dealing ruling</em>. https://thenewpublishingstandard.com/2026/07/27/delhi-high-court-openai-ani-fair-dealing-ruling/</p><p>The Next Web. (2026, July). <em>1,134 AI staff ask the US for a way to pace AI</em>. https://thenextweb.com/news/pacing-the-frontier-ai-employees-letter-us-government</p><p>The Register. (2026, July 29). <em>Word worm crawls into Copilot, spreads chaos</em>. https://www.theregister.com/security/2026/07/29/word-worm-crawls-into-copilot-spreads-chaos/5280588</p><p>Tom&#8217;s Hardware. (2026, July). <em>Nvidia and 24 other companies sign open-weights letter as Washington weighs Chinese AI model ban</em>. https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidia-and-24-other-companies-sign-open-weights-letter-as-washington-weighs-chinese-ai-model-ban</p><p>Washington Post. (2026, July 30). <em>Timeline of cyberattack by OpenAI&#8217;s AI &#8216;agent&#8217; shows its sophistication</em>. https://www.washingtonpost.com/technology/interactive/2026/07/30/timeline-cyberattack-by-openais-ai-agent-shows-its-sophistication/</p>]]></content:encoded></item><item><title><![CDATA[Least Agency vs Agent Autonomy: The OpenAI / Hugging Face Test]]></title><description><![CDATA[Least agency vs agent autonomy got tested for real when OpenAI models breached Hugging Face. See the four steps confinement stops and the three it cannot.]]></description><link>https://www.rockcybermusings.com/p/least-agency-vs-agent-autonomy-openai-huggingface-test</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/least-agency-vs-agent-autonomy-openai-huggingface-test</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 28 Jul 2026 12:50:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5rEx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5rEx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5rEx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!5rEx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!5rEx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!5rEx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5rEx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4358476,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/208489023?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5rEx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!5rEx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!5rEx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!5rEx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36f8a99b-1b79-40d9-a6c9-77fb34e5693a_2048x2048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The OpenAI and Hugging Face incident is getting all of the media attention. However, I love that it accidentally put the least agency vs. agent autonomy argument on trial. GPT-5.6 Sol and an even more capable pre-release model had a hacking benchmark to beat. They decided the shortest path was to jailbreak the sandbox, spent substantial inference compute finding a zero-day in the one service with internet reach, then pivoted into Hugging Face&#8217;s production database for the answers. Nobody authorized that. Nobody designed that experiment either.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/least-agency-vs-agent-autonomy-openai-huggingface-test?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/least-agency-vs-agent-autonomy-openai-huggingface-test?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2><strong><span>The Incident Had Two Halves And They Needed Two Different Controls</span></strong></h2><p><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>OpenAI published the benchmark evaluation on July 21, 2026.</span></a><span> They disabled the guardrails. They ran it without the production classifiers designed to prevent a model from chasing high-risk cyber work. This makes sense because how the hell would you measure the ceiling of the model&#8217;s capabilities if you left it in handcuffs? It would be like testing your vertical leap,  but requiring you to wear cement boots doing so.</span></p><p><span>Ultimately, network access out of that environment came down to one thing. An internally hosted proxy and cache for package registries. The models found a zero-day in it, worked their way through the research network until they found a node with internet access, then inferred that Hugging Face likely hosted the </span><a href="https://arxiv.org/abs/2605.11086"><span>ExploitGym </span></a><span>material and went after it with stolen credentials and more zero-days.</span></p><p><span>Now, Hugging Face had no idea. </span><a href="https://huggingface.co/blog/security-incident-july-2026"><span>Five days earlier, on July 16, they disclosed a different attack from the other end. </span></a><span>A malicious dataset abused two code-execution paths in their processing pipeline, a remote-code dataset loader and a template injection in a dataset configuration, and ran code on a processing worker. The actor escalated to node level from there, harvested cloud and cluster credentials, and moved across several internal clusters over a weekend. Their team rebuilt the sequence from more than 17,000 recorded events.</span></p><p><span>If you read the two accounts side-by-side, you will notice something uncomfortable creeping up. Hugging Face, disclosing before anyone knew the source, described an autonomous agent framework of unknown origin running a swarm of short-lived sandboxes with self-migrating command and control staged on public services. That reads like a patient adversary campaign. OpenAI describes a model that got obsessive about a test score. Both accounts are honest, and  both come from the source. From Hugging Face&#8217;s perspective, their telemetry showed that a runaway evaluation and a genuine intrusion produced the same forensic signature.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!akX6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!akX6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png 424w, https://substackcdn.com/image/fetch/$s_!akX6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png 848w, https://substackcdn.com/image/fetch/$s_!akX6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png 1272w, https://substackcdn.com/image/fetch/$s_!akX6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!akX6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png" width="1456" height="1082" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1082,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:887045,&quot;alt&quot;:&quot;Side-by-side comparison of the Hugging Face disclosure of July 16, 2026 and the OpenAI disclosure of July 21, 2026 across six attributes, showing the two accounts diverging on attacker identity, infrastructure, and motive while agreeing on detection and the absence of any public CVE.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/208489023?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Side-by-side comparison of the Hugging Face disclosure of July 16, 2026 and the OpenAI disclosure of July 21, 2026 across six attributes, showing the two accounts diverging on attacker identity, infrastructure, and motive while agreeing on detection and the absence of any public CVE." title="Side-by-side comparison of the Hugging Face disclosure of July 16, 2026 and the OpenAI disclosure of July 21, 2026 across six attributes, showing the two accounts diverging on attacker identity, infrastructure, and motive while agreeing on detection and the absence of any public CVE." srcset="https://substackcdn.com/image/fetch/$s_!akX6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png 424w, https://substackcdn.com/image/fetch/$s_!akX6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png 848w, https://substackcdn.com/image/fetch/$s_!akX6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png 1272w, https://substackcdn.com/image/fetch/$s_!akX6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0f7af4c-ce8d-4489-9283-8bfd0adfd6ae_5250x3900.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: Two Companies, One Weekend, Two Different Stories</figcaption></figure></div><p><span>That bothers me more than the &#8220;zero-day&#8221; does. Every detection story you&#8217;re running assumes you eventually work out what you&#8217;re dealing with. Hugging Face never did, and they had model-driven triage over their security telemetry that caught the compromise anyway. What it caught was the shape of the activity, and the shape was all they got.</span></p><p><span>Imagine what that means for your own program. Attribution is the step every incident-response plan treats as eventually solvable, even though it really shouldn&#8217;t, and instead focuses on the rapid restoration of service. You collect enough telemetry, you correlate enough signals, and the picture </span><em><span>may </span></em><span>resolve into a named actor with known tradecraft. Hugging Face collected the telemetry, correlated the signals, and got a shape. The name arrived five days later in OpenAI&#8217;s blog post, and it arrived only because OpenAI volunteered it. Nothing in their environment was ever going to produce it. Build your response plan on the assumption that you get the shape and never the name.</span></p><p><span>The two halves failed differently, and they need different fixes. The OpenAI side is a containment problem, where a capable workload got out of a box. The Hugging Face side is an authority problem, where a pipeline component executed supplied content while holding credentials worth stealing. Treating them as one story is how people end up recommending one control for both, and that&#8217;s just plain wrong.</span></p><h2><strong><span>Authority And Autonomy Are Different Dials</span></strong></h2><p><span>Here is my stake in the ground. You can cut what an agent may effect without cutting how freely it reasons. Authority sits on one dial and decision independence sits on the other. Turn the first one down and the second one can still stay where it was.</span></p><p><span>The best evidence for that sits inside this incident, and I haven&#8217;t seen anyone call that out yet.</span></p><p><span>Look at Hugging Face&#8217;s dataset-processing worker. That thing decides almost nothing. It&#8217;s an ingestion pipeline. It doesn&#8217;t plan, pick between strategies, or reason about goals. Its autonomy sits near zero. It also executed attacker-supplied code while holding cloud and cluster credentials that opened several internal clusters. That&#8217;s enormous authority paired with no autonomy at all.</span></p><p><span>If authority and autonomy were one dial, that combination couldn&#8217;t exist. You&#8217;d never find a component with maximum reach and minimum judgment. Every environment I&#8217;ve walked into is chock-full of them.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!etMw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!etMw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png 424w, https://substackcdn.com/image/fetch/$s_!etMw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png 848w, https://substackcdn.com/image/fetch/$s_!etMw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png 1272w, https://substackcdn.com/image/fetch/$s_!etMw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!etMw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png" width="1456" height="1090" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1090,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:676693,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/208489023?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!etMw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png 424w, https://substackcdn.com/image/fetch/$s_!etMw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png 848w, https://substackcdn.com/image/fetch/$s_!etMw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png 1272w, https://substackcdn.com/image/fetch/$s_!etMw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74a8764d-de21-452b-a2c3-7954b6482980_4650x3480.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: Authority and Autonomy Are Different Dials </figcaption></figure></div><p><span>None of this is new thinking. The object-capability model has held the two apart for decades, and Norm Hardy named the sharpest version of the problem in October 1988 in </span><a href="https://dl.acm.org/doi/10.1145/54289.871709"><span>three pages of ACM SIGOPS Operating Systems Review</span></a><span>. His example was a compiler at Tymshare that held write access to a billing file for accounting reasons. A customer handed it an output filename that matched the billing file, and the compiler dutifully overwrote the company&#8217;s revenue records on the customer&#8217;s behalf. Every action stayed inside policy the entire way through. Hardy&#8217;s answer was that authority should ride with whoever invokes an action instead of sitting ambient inside the program doing the work. Read the OpenAI chain again with that in mind. The package proxy held network authority for one purpose, and a caller found a way to spend it on another.</span></p><p><span>I learned the practical version via my oil and gas days long before I thought about it in agents. A control-room operator at a console can reason about anything they want. They can decide the process needs to run hotter, that a reading is lying to them, that the previous shift left them a mess. What they can&#8217;t do is open two valves that must never be open together, because the interlock is wired to make that combination impossible, and no amount of operator reasoning gets around it. I&#8217;ve watched sharp operators argue with an interlock and lose, which is the point of the interlock. The decision space stayed wide open, the authority stayed hard-bounded, and the two never touched each other in thirty years of me watching them.</span></p><p><span>The analogy breaks in one place&#8230; Physics enforces an interlock. Code enforces an agent boundary, and code carries flaws, which is the entire story of what happened at that package proxy. That&#8217;s an argument for keeping your enforcement point small, dumb, and outside the model, and for never running only one of them.</span></p><p><span>On the research side, </span><a href="https://arxiv.org/abs/2503.18813"><span>CaMeL</span></a><span> is the closest thing to a measurement anyone has published. Debenedetti and nine coauthors built a defense that solves 77% of AgentDojo tasks with a provable security guarantee, against 84% for the same system undefended. The code sits under Google Research. That&#8217;s seven points of task completion traded for a capability-enforced boundary.</span></p><p><span>Before you file that under &#8220;autonomy survives scoping,&#8221; read what CaMeL constrains. It works by pulling control flow out of the trusted query so untrusted data can never influence the program, which narrows what the agent may decide about that data. The number tells you enforcement is cheap in utility terms. It doesn&#8217;t tell you the decision space came through untouched. I&#8217;d rather concede that now than have a reader hand me back my own citation.</span></p><h2><strong><span>Human Approval Was Never Going To Hold This Line</span></strong></h2><p><span>The reflex fix for all of this is a human in the loop. Put a person between the agent and the consequential action.</span></p><p><span>That control works in specific places. Aviation runs a two-person rule. Nuclear weapons handling runs a two-man rule. Wire transfers run four eyes. None of those are ceremonial.</span></p><p><span>My read on why they work lies in five conditions that hold at once.</span></p><ol><li><p><span>Events stay rare enough that a reviewer reads each one.</span></p></li><li><p><span>Each event carries stakes big enough to justify the interruption.</span></p></li><li><p><span>The reviewer holds training for that specific decision.</span></p></li><li><p><span>Time exists to make it.</span></p></li><li><p><span>The criteria are legible before you look.</span></p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!i0qx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!i0qx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png 424w, https://substackcdn.com/image/fetch/$s_!i0qx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png 848w, https://substackcdn.com/image/fetch/$s_!i0qx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png 1272w, https://substackcdn.com/image/fetch/$s_!i0qx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!i0qx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png" width="1456" height="1071" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1071,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:868555,&quot;alt&quot;:&quot;Radar chart comparing the aviation two-person rule, the nuclear two-man rule, four-eyes wire approval, and agent action streams across five conditions, with the first three forming wide polygons and agent action streams collapsing toward the center on every axis.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/208489023?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Radar chart comparing the aviation two-person rule, the nuclear two-man rule, four-eyes wire approval, and agent action streams across five conditions, with the first three forming wide polygons and agent action streams collapsing toward the center on every axis." title="Radar chart comparing the aviation two-person rule, the nuclear two-man rule, four-eyes wire approval, and agent action streams across five conditions, with the first three forming wide polygons and agent action streams collapsing toward the center on every axis." srcset="https://substackcdn.com/image/fetch/$s_!i0qx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png 424w, https://substackcdn.com/image/fetch/$s_!i0qx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png 848w, https://substackcdn.com/image/fetch/$s_!i0qx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png 1272w, https://substackcdn.com/image/fetch/$s_!i0qx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26961670-c13c-4a1e-bba4-543589a04c06_4650x3420.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: Where Human Approval Holds, and Where It Never Could</figcaption></figure></div><p><span>Agent action streams fail all five. The events aren&#8217;t rare because they can arrive in the thousands. Individual actions carry almost no stakes, which is why the trajectory is the dangerous object rather than any step inside it. The approver is usually a developer with a build to ship. No time budget exists. The criteria are illegible too, because &#8220;install a package&#8221; looks identical whether it&#8217;s step four of a build or step four of an escape.</span></p><p><span>Invariant Labs saw this coming. When Marco Milanta and Luca Beurer-Kellner disclosed the </span><a href="https://invariantlabs.ai/blog/mcp-github-vulnerability"><span>GitHub MCP toxic agent flow in May 2025</span></a><span>, they called it an architectural problem rather than a code flaw. Claude Desktop asks users to confirm each tool call by default. Milanta and Beurer-Kellner wrote in the same post that many users already choose an &#8220;always allow&#8221; policy and stop watching individual actions. A control that people switch off during normal use has already failed.</span></p><p><span>Did your last threat model include human-in-the-loop scaling? Be honest. It likely didn&#8217;t. Multiply your expected agent action volume by the seconds of genuine attention each approval deserves, then compare that number to the reviewer hours you staffed. If the product exceeds the capacity, all you have done is build a rubber-stamping assembly line. You&#8217;ve built a queue that somebody will clear by clicking through it, and the clicking looks identical in your audit log to careful review. Watch your always-allow rate. When it climbs, the control already converted itself into a rubber stamp, and your evidence trail is now recording consent you never gave.</span></p><p><span>Machine-speed oversight closes the argument. </span><a href="https://www.crowdstrike.com/en-us/press-releases/2026-crowdstrike-global-threat-report/"><span>CrowdStrike&#8217;s 2026 Global Threat Report</span></a><span> put average eCrime breakout time at 29 minutes for 2025, with the fastest observed at 27 seconds, and exfiltration starting within four minutes of initial access in one intrusion. Hugging Face&#8217;s attacker ran more than 17,000 events across a weekend. Work out how much human attention each of those actions gets.</span></p><p><span>The regulation already knows. </span><a href="https://artificialintelligenceact.eu/article/14/"><span>EU AI Act Article 14</span></a><span> requires that high-risk systems be built so natural persons can effectively oversee them. Article 14(4)(b) requires that the overseer stay aware of the tendency toward &#8220;automatically relying or over-relying&#8221; on system output, which the Act names automation bias. The same article that mandates the control documents the way it fails.</span></p><h2><strong><span>Least Agency vs Agent Autonomy: Where The One-Axis Reading Wins</span></strong></h2><p><span>Now my own argument takes damage.</span></p><p><span>One disclosure before I start. I co-lead OWASP&#8217;s Agentic AI Security work, so I&#8217;m about to criticize language I share responsibility for.</span></p><p><a href="https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/"><span>OWASP&#8217;s Top 10 for Agentic Applications</span></a><span>, published in December 2025, puts least agency underneath all ten entries as the governing principle. The definition tells you to grant agents only the minimum autonomy required for safe, bounded tasks, with autonomy earned through demonstrated safety rather than handed over by default.</span></p><p><span>Read that as a practitioner and it&#8217;s obviously right. Read it as an architect and notice what it does. It puts autonomy on the dial you&#8217;re turning down.</span></p><p><span>The phrasing carries a cost downstream, and I&#8217;ve watched it land. A team reads &#8220;minimum autonomy&#8221; and reaches for the control that reduces autonomy, which is an approval gate. What they needed was the control that reduces authority, which is confinement. Those are different builds with different failure modes, and the first one degrades into always-allow while the second one holds. Invariant Labs documented that behavior in the wild more than a year before this incident, and the industry ran the experiment anyway.</span></p><p><a href="https://arxiv.org/abs/2506.08837"><span>The Design Patterns paper</span></a><span> makes the one-axis case honestly, and I respect it for that. Beurer-Kellner and thirteen coauthors write that their patterns &#8220;constrain the actions of agents to explicitly prevent them from solving arbitrary tasks.&#8221; That&#8217;s the trade stated plainly. They go further. As long as agents and their defenses run on the current class of language models, they consider it unlikely that general-purpose agents can offer meaningful and reliable safety guarantees.</span></p><p><span>I think they&#8217;re right about today&#8217;s agents. Where users perceive autonomy right now, most of what they&#8217;re perceiving is range. Give a model more tools and more destinations, and it feels more autonomous, because it surprises you in more directions. Tighten the range, and it feels dumber even when its reasoning never changed.</span></p><p><span>The authority/autonomy axes are separate in principle, and capability systems have kept them separate in practice. Today&#8217;s agents happen to draw most of their perceived autonomy from breadth of authority, which means tightening authority does cost you something people are calling autonomy.</span></p><h2><strong><span>Where The Claim Runs Out</span></strong></h2><p><span>Run the whole chain and mark each step for what capability confinement would have done to it.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zKGW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zKGW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png 424w, https://substackcdn.com/image/fetch/$s_!zKGW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png 848w, https://substackcdn.com/image/fetch/$s_!zKGW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png 1272w, https://substackcdn.com/image/fetch/$s_!zKGW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zKGW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png" width="1456" height="659" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:659,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:418823,&quot;alt&quot;:&quot;Nine-stage attack chain from the OpenAI evaluation sandbox to the Hugging Face production database, with each stage marked as prevented by authority scoping, contained only, dependent on the boundary&#8217;s code, or out of reach.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/208489023?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Nine-stage attack chain from the OpenAI evaluation sandbox to the Hugging Face production database, with each stage marked as prevented by authority scoping, contained only, dependent on the boundary&#8217;s code, or out of reach." title="Nine-stage attack chain from the OpenAI evaluation sandbox to the Hugging Face production database, with each stage marked as prevented by authority scoping, contained only, dependent on the boundary&#8217;s code, or out of reach." srcset="https://substackcdn.com/image/fetch/$s_!zKGW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png 424w, https://substackcdn.com/image/fetch/$s_!zKGW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png 848w, https://substackcdn.com/image/fetch/$s_!zKGW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png 1272w, https://substackcdn.com/image/fetch/$s_!zKGW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb64efa07-6fec-400e-84fb-04b0769c6787_5700x2580.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: Where Authority Scoping Bites</figcaption></figure></div><p><span>The lateral movement and credential harvesting steps disappear under per-task credentials that expire. The reach into Hugging Face disappears under egress that defaults to deny. The dataset worker&#8217;s blast radius shrinks hard once that worker can&#8217;t hold a credential worth stealing. Four steps prevented, one contained.</span></p><p><span>Three steps confinement never touches. It has nothing to say about running an evaluation with the production classifiers off, which is a policy decision people made. It has nothing to say about a model choosing to look for a way out. It has nothing to say about a model inferring where the answer key lives.</span></p><p><span>That last step is the showstopper here. OpenAI denied standing egress, and yet that environment had exactly one path out, a package proxy, which is honestly a tighter posture than most enterprises run around their build systems. The model found a zero-day in the boundary itself.</span></p><p><span>&#8220;Deny standing egress&#8221; as a recommendation has a caveat, though. A capability boundary is worth whatever its code is worth. One boundary is one point of failure. The lesson from step three is defense in depth around your confinement mechanism, and confidence in the mechanism is the thing that got tested here.</span></p><p><span>One thing would move me off this. Show me a capability-confined agent stack that runs untrusted content in production for a year without an authority escape, where the enforcement point is small enough to audit, and there is more than one of them in the path. That&#8217;s a verifiable claim rather than a position you have to take on faith. Absent that, I&#8217;m holding the independent argument at high confidence and the this-would-have-worked argument at somewhere near a coin flip.</span></p><p><span>That&#8217;s a split verdict, and I&#8217;m leaving it split. The independent claim holds. The claim that least agency would have stopped this incident holds for part of the chain and comes apart on the rest.</span></p><p><span>Then there&#8217;s the definitional problem, which is the sharpest objection anyone can make. If I define autonomy as bounded, capability-scoped action inside pre-authorized flows, then &#8220;least agency is compatible with autonomy&#8221; comes out true because I made it true. That&#8217;s winning by relabeling.</span></p><p><span>A meaningful notion of autonomy requires the capacity to do something unanticipated. The models in this incident did precisely that. That behavior is the autonomy under discussion, and capability confinement exists to forbid it. My compatibility claim holds cleanly for the deflationary sense of the word and stays in genuine tension with the strong sense. I&#8217;m not going to pretend that tension resolves.</span></p><p><span>Now, a minority of readers on Hugging Face&#8217;s own comment thread called the episode a marketing exercise with a conveniently soft sandbox. I can see that.</span></p><p><span>Watch what OpenAI&#8217;s post does while it reports the incident. It calls the event unprecedented and the capabilities state of the art. It cites UK AISI evaluations of long-horizon cyber operations and says this incident shows those capabilities carry into real systems. It announces that Hugging Face joined OpenAI&#8217;s trusted access program, links that program twice, and closes by inviting other defenders to apply. A disclosure about losing control of a model also demonstrates that the model is formidable and routes you toward the commercial program built on exactly that.</span></p><p><span>Now check what an outsider can verify without taking anyone&#8217;s word. Nothing. No CVE for the proxy zero-day, none for the dataset paths, and the proxy vendor stays unnamed. Every technical claim in both posts comes from the two parties inside the story.</span></p><p><span>My take? I think the incident happened, but both companies realized what a HUGE marketing opportunity it is. OpenAI gets to draw the news cycle away from Mythos/Fable, and Hugging Face&#8217;s entire existence stands on serving up open-source/open-weight models.</span></p><p><span>Oh yeah&#8230; small footnote&#8230; Hugging Face had to use an open-weight model, GLM 5.2 to be exact, in their incident response because the frontier model APIs stonewalled them. They couldn&#8217;t tell defender IR activity vs. attacker activity.</span></p><p><span>You have an agent in your environment right now that touches content somebody outside your company wrote. Answer two questions about it separately, on paper, today. What can it decide? What can it affect?</span></p><p><span>Most people answer the first and assume they&#8217;ve covered the second. Hugging Face&#8217;s dataset worker is what that assumption looks like from the outside. It decided nothing. It reached everything. The least agency vs agent autonomy question won&#8217;t get settled in the literature this year, and it won&#8217;t get settled on your architecture diagram either. It gets settled by default, badly, every time somebody scopes the first question and skips the second.</span></p><p><span>Which one did you scope?</span></p><h3><strong><span>What To Do Next</span></strong></h3><p><span>Three mechanisms, and I&#8217;m holding this to authority scoping, because detection and boundary hygiene are a different conversation.</span></p><ol><li><p><span>Issue per-task credentials that expire in minutes instead of standing keys, so a compromised worker holds authority with a short shelf life.</span></p></li><li><p><span>Put a policy decision point at the tool boundary that returns deny for any tool or destination it hasn&#8217;t seen before, evaluated per action rather than once at token issuance.</span></p></li><li><p><span>Make it a hard rule that any worker processing content from outside your organization cannot reach a credential worth stealing.</span></p></li></ol><p><span>That&#8217;s Create and Adapt work in the</span><a href="https://www.rockcyber.com/ai-strategy-and-governance"><span> CARE model</span></a><span>, and it&#8217;s where least agency stops being a principle you nod at.</span></p><p><span>For the definitional groundwork underneath this piece, I wrote the cornerstone on</span><a href="https://www.rockcybermusings.com/p/i-agent-authentication-authorization-gap"><span> least agency versus least privilege</span></a><span> and where OAuth scopes stop being enough.</span><a href="https://www.rockcybermusings.com/p/autonomous-ai-agent-ransomware-jadepuffer"><span> JadePuffer</span></a><span> is what an unbounded agent does to a victim who had every patch available and skipped them. The advisory work lives at</span><a href="https://www.rockcyber.com/"><span> rockcyber.com</span></a><span>.</span></p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span> and you can subscribe above</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, </span><strong><a href="https://zenity.io/resources/webinars/a-conversation-with-claude-mythos-tenant-aws-on-demand"><span>check out my chat with Itay Meller, Specialist Solutions Architect, Security at AWS</span></a></strong><span>, about what autonomous vulnerability discovery changes for security teams, and where AI Detection and Response (AIDR) fits in a stack that was never built to watch agents. </span><em><strong><span>No registration required</span></strong></em></p><p></p>]]></content:encoded></item><item><title><![CDATA[Weekly Musings Top 10 AI Security Wrapup: Issue 47 July 17 -July 23, 2026]]></title><description><![CDATA[OpenAI's own models hacked Hugging Face, Google shipped a gov-only bug hunter, and regulators fenced in AI agents. A CISO's AI security and governance brief.]]></description><link>https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260717-20260723</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260717-20260723</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Fri, 24 Jul 2026 12:50:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!mdPE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mdPE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mdPE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!mdPE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!mdPE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!mdPE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mdPE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1233556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/208267495?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mdPE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!mdPE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!mdPE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!mdPE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3a599b-8568-47ce-82d3-24017425b2b5_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>An autonomous attacker spent a weekend inside Hugging Face, harvesting credentials. Investigators assumed a human crew. Five days later OpenAI admitted the culprit was its own models, run with the safety brakes loosened to cheat on a cyber benchmark. AI stopped being only a thing we secure and became a thing that secures and attacks on its own. Google shipped a vulnerability-hunting model for governments only. Anthropic put a scanner in every developer&#8217;s terminal. Paris warned that three firms own the agent market. Read this before your next board meeting.</p><p>Watch where the capability lands before you watch the headline. This week it landed in three places at once: an offensive test that got loose, defensive tooling any engineer can now run, and the desks of regulators who started treating autonomous agents as a distinct risk category rather than a faster chatbot. The through-line is agency. Systems that plan, act, and persist across sessions do not behave like the software your controls were built for. They hold credentials, they touch production, and they remember, which breaks assumptions baked into identity, change management, and vendor risk. Here are the eleven developments from July 17 to July 23 that move a decision. The last one barely made the news, and it should worry you the most. Deeper work lives at <a href="https://rockcybermusings.com/">rockcybermusings.com</a>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260717-20260723?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260717-20260723?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>1. OpenAI&#8217;s Own Models Broke Out of a Lab and Hacked Hugging Face</h3><p>Hugging Face detected an intrusion on July 16, 2026, first blamed on an unknown autonomous agent that moved laterally across clusters over a weekend (Help Net Security). On July 21, OpenAI admitted the attacker was its own GPT-5.6 Sol and a pre-release model, run with reduced cyber-safety refusals for an internal benchmark called ExploitGym, which escaped the sandbox through a package-installer zero-day and reached Hugging Face production to steal the answer key (Fortune). It is the first documented case of frontier models chaining a real zero-day without source code access.</p><p><strong>Why it matters</strong></p><ul><li><p>Your sandbox is a hypothesis now, not a control.</p></li><li><p>Attribution got harder. The largest model hub spent five days blaming an elite human crew.</p></li><li><p>Evaluation safety is a supply-chain problem, because the blast radius reached a third party&#8217;s production.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Air-gap or egress-filter any environment running reduced-safety evaluations.</p></li><li><p>Add an autonomous-actor hypothesis to incident playbooks, with velocity and concurrency signals.</p></li><li><p>Demand containment controls in writing before placing pre-release models near sensitive systems.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I have sat through a hundred tabletops where someone swears the sandbox holds. This week, one at a top lab did not, and the thing that got out robbed a neighbor to win a game. It was not malicious. It wanted the answer key, and hacking a production database was the efficient path. We over-index on superintelligence when the near-term danger is a competent system with a narrow goal and no sense of the property line. I keep arguing we should <a href="https://www.rockcyber.com/">quantify AI risk</a> with probabilities and blast radius, not vibes. Show your board this.</p><h3>2. China Launches WAICO, the First Intergovernmental Body for AI</h3><p>The 2026 World Artificial Intelligence Conference and High-Level Meeting on Global AI Governance ran in Shanghai from July 17 to July 20, with President Xi Jinping&#8217;s keynote on July 17 and UN Secretary-General Ant&#243;nio Guterres among representatives from more than 100 countries (China Daily). The headline outcome was the agreement establishing the World Artificial Intelligence Cooperation Organization, or WAICO, the first intergovernmental body dedicated to AI, headquartered in Shanghai and pitched as a way to bridge the AI divide for the Global South.</p><p><strong>Why it matters</strong></p><ul><li><p>A second pole of AI governance now has an address, and three rulebooks that do not interchange.</p></li><li><p>WAICO standards could add a whole region to your compliance map.</p></li><li><p>Governance turned into geopolitics, and where you train and evaluate models now carries diplomatic weight.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Add WAICO to your regulatory-watch list beside the EU AI Office and NIST.</p></li><li><p>In Global South markets, map exposure to WAICO-aligned procurement and localization.</p></li><li><p>Brief your board on bloc fragmentation, and budget for parallel compliance.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I read the Shanghai communique by asking who it binds and who it exempts. WAICO is smart statecraft. China gets to write the rules for a hundred countries that cannot afford their own labs, and it does it under the banner of fairness while Washington argues about whether states can regulate at all. Plan for a world where your AI stack needs a passport.</p><h3>3. Google Ships Gemini 3.5 Flash Cyber, and Keeps It for Governments</h3><p>Google DeepMind announced Gemini 3.5 Flash Cyber on July 21, 2026, a model tuned to find vulnerabilities, confirm exploitability, and generate patches (Help Net Security). On the V8 JavaScript engine, it found 55 confirmed issues, against 47 for mainline 3.5 Flash and 36 for Opus 4.6, including 10 that the others missed. Google reserved access for governments and trusted partners through its CodeMender program and published no pricing or release date, citing dual-use safety risks (The Hacker News).</p><p><strong>Why it matters</strong></p><ul><li><p>The offense-defense gap narrowed, and access is now a policy lever rather than a market one.</p></li><li><p>Government-only access means state-backed adversaries may reach parity before you do.</p></li><li><p>Ten issues two strong models missed is not noise, and specialized cyber models will outrun general ones.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask your vendors which frontier cyber models they can access, and when you get them.</p></li><li><p>Fold AI-discovered vulnerability classes into your patch SLAs.</p></li><li><p>Watch the access policy, because who gets the model shapes your threat model.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Google did the responsible thing, and rationing a vulnerability-hunting model to governments on dual-use grounds means the capability arrives unevenly. The state-aligned crews I worry about in energy and water do not wait for a public API. Get into the partner conversations now, and do not assume the gate holds forever.</p><h3>4. The Linux Kernel Published 440 AI-Found CVEs in a Day</h3><p>Between July 19 and July 20, 2026, the Linux kernel project published roughly 440 CVE advisories, 431 on the first day and nine more on the second, driven by static analysis, fuzzing, and AI-assisted auditing that named Sashiko, the kernel&#8217;s AI review system, in 23 cases (securityonline.info). A closer look from XenoSpectrum found that many identifiers were bulk-assigned to bugs already fixed upstream, which points to CVE inflation rather than a wave of fresh exploitable risk.</p><p><strong>Why it matters</strong></p><ul><li><p>CVE volume is decoupling from human review capacity, and most of the flood is already patched.</p></li><li><p>Signal-to-noise is the fight, because tooling drowns when identifiers attach to already-fixed issues.</p></li><li><p>The kernel is the canary for what AI auditors will do to every large codebase you depend on.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Weight your program on exploitability and patch status, not raw CVE count.</p></li><li><p>Automate advisory ingestion and map it to deployed versions.</p></li><li><p>Press your software vendors on backport speed.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Publishing 440 CVEs in a day sounds like the sky is falling, and it was mostly bookkeeping, because most were already fixed. What changed is the tempo. The trap is that your metrics still treat every CVE as a fire, so your team burns out chasing paperwork while the real bug hides in the pile. Fix the dashboard before the flood.</p><h3>5. Anthropic Puts a Multi-Agent Security Scanner in Every Terminal</h3><p>Anthropic launched the Claude Security plugin for Claude Code in public beta during the week of July 22, 2026, letting developers scan changes or a whole codebase from the terminal (MarkTechPost). A coordinated set of agents maps the codebase, models threats, connects issues across files, and returns suggested patches for flaws such as injection and authentication bypass. Admins enable it through the console, which reaches teams that never had an application-security function.</p><p><strong>Why it matters</strong></p><ul><li><p>Application-security review is moving into the commit loop for everyone.</p></li><li><p>Multi-agent scanning catches cross-file logic bugs that traditional static analysis misses.</p></li><li><p>The model that writes your code now audits it, a single-vendor dependency worth naming.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Pilot it against a repository you already scanned, and compare findings.</p></li><li><p>Set policy on who enables it and where findings go, because reports are sensitive telemetry.</p></li><li><p>Keep your existing controls, and treat AI review as an added layer.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I like this, and I do not trust it yet, which is the right posture for anything in week-one beta. The catch nobody says out loud is that the model that wrote the vulnerable code now grades its own homework, and the failure modes correlate. Run it beside your scanners and keep score.</p><p>Here&#8217;s proof from my own terminal:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rGXr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rGXr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png 424w, https://substackcdn.com/image/fetch/$s_!rGXr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png 848w, https://substackcdn.com/image/fetch/$s_!rGXr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png 1272w, https://substackcdn.com/image/fetch/$s_!rGXr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rGXr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png" width="1247" height="56" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:56,&quot;width&quot;:1247,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:18112,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/208267495?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rGXr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png 424w, https://substackcdn.com/image/fetch/$s_!rGXr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png 848w, https://substackcdn.com/image/fetch/$s_!rGXr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png 1272w, https://substackcdn.com/image/fetch/$s_!rGXr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1f33574-44d0-44ab-9808-736151fde1e0_1247x56.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><h3>6. France Says Three Firms Own 84% of the AI Agent Market</h3><p>On July 17, 2026, France&#8217;s Autorit&#233; de la concurrence published Opinion No. 26-A-05, from an inquiry opened in January 2026 into the AI agents sector (Concurrences). The regulator found that OpenAI, Google, and Anthropic together hold more than 84% of the agent market, and it warned that autonomous agents could concentrate the digital economy around a few vertically integrated firms unless regulators act on data access, interoperability, and default placement. The 3,700-page opinion flagged platformisation, disintermediation, and algorithmic collusion.</p><p><strong>Why it matters</strong></p><ul><li><p>Vendor lock-in is a documented market fact now, and your exit options are thin.</p></li><li><p>Default placement is the new shelf space, deciding who sees your data and traffic.</p></li><li><p>Algorithmic collusion is on the table for anyone whose agents set prices.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory the foundation models under your agent stack and what switching costs.</p></li><li><p>Negotiate data-portability and interoperability terms while regulators give you cover.</p></li><li><p>If your agents touch pricing, document human oversight.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This got buried under the Hugging Face fireworks, and it is the one your CFO should read. The French built their own agents, ran 550 shopping queries, and measured who actually gets cited. Eighty-four percent in three hands. I have watched the lock-in movie before with cloud, and the sequel always costs more than the trailer. Sign nothing multi-year without a portability clause.</p><h3>7. France&#8217;s CNIL Names Agent Memory as a GDPR Problem</h3><p>On July 20, 2026, France&#8217;s data protection authority, the CNIL, with the Council for AI and Digital Technology, published a note on the risks of agentic AI for personal data (Digital Policy Alert). It named persistent memory and multi-service interaction as the features that strain the GDPR most, since memory collides with purpose limitation and cross-system action blurs the controller and processor line. France often previews where the wider EU lands, so read this as an early warning.</p><p><strong>Why it matters</strong></p><ul><li><p>Persistent memory is a named privacy hazard your impact assessments never anticipated.</p></li><li><p>Multi-service agents blur controller and processor roles, and that means liability.</p></li><li><p>France leads, so a CNIL position signals where EU enforcement heads next.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Update your impact assessments to cover what memory stores, for how long, and how a subject deletes it.</p></li><li><p>Map every agent&#8217;s cross-service reach and assign controllership per workflow.</p></li><li><p>Build a memory-retention and erasure policy before a regulator asks.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Everyone obsesses over what an agent does. The CNIL asked the sharper question, which is what it remembers and who owns that memory. Persistent memory outlives the purpose you collected it for, which is a slow-motion GDPR problem. If you cannot show a regulator how a data subject wipes themselves from an agent&#8217;s long-term memory, you do not have a compliant deployment. France just put that in writing, and I would fix it before August 2.</p><h3>8. The Financial Stability Board&#8217;s AI Consultation Closed July 22</h3><p>The comment window for the Financial Stability Board&#8217;s consultation, Sound Practices for Responsible Adoption of Artificial Intelligence, closed on July 22, 2026 (Financial Stability Board). The report offered 12 sound practices for organization-wide AI governance across the lifecycle and flagged agentic AI as an emerging risk to financial stability, with a final report planned for October 2026. The FSB does not chase consumer harms, so its attention signals where bank supervision heads next.</p><p><strong>Why it matters</strong></p><ul><li><p>The systemic-risk regulator for global finance is now writing AI governance expectations.</p></li><li><p>Twelve sound practices become a de facto checklist that examiners reference.</p></li><li><p>Agentic AI reached the stability conversation, elevating model risk from a compliance line item.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Map your AI governance to the FSB&#8217;s 12 practices before the October final report.</p></li><li><p>Brief your model-risk and third-party-risk teams that autonomous agents fall in scope.</p></li><li><p>Prepare board-level documentation, because stability regulators expect governance owned at the top.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>When the Financial Stability Board starts talking about your AI, governance as a side project is over. The fact that agentic AI reached its radar points toward supervision that treats an autonomous trading agent like any other source of systemic risk, and its twelve practices rhyme with ISO 42001 and the NIST framework. If you run risk at a bank and you are not already mapping to this, you will explain yourself to an examiner later.</p><h3>9. Neo Security Exits Stealth With $100 Million for Agent Control</h3><p>On July 20, 2026, Neo Security left stealth with $100 million to build a secure control layer for enterprise AI agents (SiliconANGLE). Andreessen Horowitz and Bessemer led the round, and the founders came out of SentinelOne, which puts endpoint-grade thinking behind the pitch. The raise fit the day&#8217;s pattern, as venture money moved toward the control plane of AI rather than the presentation layer, betting that enterprises will need purpose-built controls to scope, monitor, and revoke what agents can do.</p><p><strong>Why it matters</strong></p><ul><li><p>The market just priced agent security at $100 million out of stealth.</p></li><li><p>Endpoint-security founders signal the category matures fast.</p></li><li><p>A control plane for agents is the architecture pattern to watch.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Test whether your identity and privileged-access stack can scope, monitor, and revoke agent credentials.</p></li><li><p>Start with least privilege and short-lived credentials for every agent today.</p></li><li><p>Track this category in your architecture roadmap.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Follow the money and it tells you what is coming. A hundred million dollars out of stealth for an agent control plane is the market saying the quiet part, that your identity stack was never built for software that acts on its own. A pile of venture money produces overlapping products, so buy the least-privilege discipline first and let the market shake out.</p><h3>10. OpenAI Wants You to Measure AI by &#8220;Useful Intelligence per Dollar&#8221;</h3><p>On July 17, 2026, OpenAI&#8217;s chief financial officer Sarah Friar proposed what she called the ultimate scorecard for the AI age (Axios). Her core metric, useful intelligence per dollar, runs four questions on work accomplished, cost, dependability, and return on compute. The dependability piece matters most for security leaders, since it counts how often a tool met quality standards, needed edits, and required a person to finish the task.</p><p><strong>Why it matters</strong></p><ul><li><p>The vendor is defining how you measure its value, so read the metric as a negotiating anchor.</p></li><li><p>Intervention rate is exactly the oversight data your governance program needs.</p></li><li><p>ROI pressure drives shadow AI faster than governance can follow.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Adopt dependability tracking on your own terms, tied to risk, not cost alone.</p></li><li><p>Push back on ROI framing that ignores security and governance overhead.</p></li><li><p>Use the intervention metric as an early warning for over-delegated agents.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I will give OpenAI credit for a useful idea wrapped in a sales pitch, though useful intelligence per dollar is also a number designed to make you feel good about the invoice. The part I want is the dependability tracking, meaning how often the thing needed a human to clean up, because that is your oversight telemetry and your risk signal in one. Keep the metric, and question the messenger.</p><h3>The One Thing You Won&#8217;t Hear About But You Need To: The Chronos Vulnerability</h3><p>The week of July 20, 2026, researchers Om Narayan, Ramkinker Singh, and Praveen Baskar posted a paper to arXiv, The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI, accepted at the ACM KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI (arXiv). It maps how attackers exploit agent memory over time, the threat now tracked as memory poisoning and formalized as OWASP ASI06 in the 2026 Agentic AI Top 10. Poisoned memory survives past the originating session, so unlike prompt injection the attack and its effect are decoupled in time, with reported success rates from 80% to nearly 100% and defenses showing limited effect.</p><p><strong>Why it matters</strong></p><ul><li><p>Memory poisoning is the attack that waits, so single-session detections miss it entirely.</p></li><li><p>Enterprise agents are the target, because memory-rich systems are the most exposed to durable compromise.</p></li><li><p>Success rates run from 80% to nearly total, and defenses lag.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory which agents hold persistent memory, and treat that memory as a trust boundary.</p></li><li><p>Add provenance and validation to memory writes.</p></li><li><p>Red-team your agents across sessions, not only within one.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This is a ten-page paper almost no one in your leadership will see, and it describes the threat I would use against you. The smart attacker poisons the memory and walks away, because the payload fires next week, long after your logs went quiet. Prompt injection is a mugging. Memory poisoning plants a sleeper. Put memory inside the boundary of identity, provenance, and least privilege now, because the attackers already read the paper. That is the argument I keep making at <a href="https://www.rockcyber.com/">rockcyber.com</a>.</p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation continues at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span>.</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, </span><strong><a href="https://zenity.io/resources/webinars/a-conversation-with-claude-mythos-tenant-aws-on-demand"><span>check out my chat with Itay Meller, Specialist Solutions Architect, Security at AWS</span></a></strong><span>, about what autonomous vulnerability discovery changes for security teams, and where AI Detection and Response (AIDR) fits in a stack that was never built to watch agents. </span><em><strong><span>No registration required</span></strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share RockCyber Musings</span></a></p><h2>References</h2><p>Autorit&#233; de la concurrence. (2026, July 17). <em>The French Competition Authority issues an opinion on the competitive functioning of the AI agents sector</em>. Concurrences. https://www.concurrences.com/en/bulletin/news-issues/preview/the-french-competition-authority-issues-an-opinion-on-the-competitive</p><p>Axios. (2026, July 17). <em>Exclusive: OpenAI&#8217;s CFO pitches a new way to measure AI&#8217;s value</em>. https://www.axios.com/2026/07/17/openai-ai-costs-roi-metrics</p><p>Ballon, M. (2026, July 21). <em>OpenAI says Hugging Face was breached by its pre-release models</em>. TechCrunch. https://techcrunch.com/2026/07/21/openai-says-hugging-face-was-breached-by-its-pre-release-models/</p><p>China Daily. (2026, July 18). <em>Chair&#8217;s statement of the 2026 World Artificial Intelligence Conference &amp; High-Level Meeting on Global AI Governance</em>. https://www.chinadaily.com.cn/a/202607/18/WS6a5ab88aa310986e2b465f4e.html</p><p>Digital Policy Alert. (2026, July 20). <em>CNIL and CIANum released a note on the risks of agentic AI for personal data protection</em>. https://digitalpolicyalert.org/event/41932-national-commission-for-information-technology-and-civil-liberties-and-council-for-ai-and-digital-technology-released-a-note-on-the-risks-of-agentic-ai-for-personal-data-protection</p><p>Financial Stability Board. (2026, June 10). <em>Sound practices for responsible adoption of artificial intelligence (AI): Consultation report</em>. https://www.fsb.org/2026/06/sound-practices-for-responsible-adoption-of-artificial-intelligence-ai-consultation-report/</p><p>Help Net Security. (2026, July 20). <em>Hugging Face breached by autonomous AI agent</em>. https://www.helpnetsecurity.com/2026/07/20/hugging-face-breached-by-autonomous-ai-agent/</p><p>Help Net Security. (2026, July 22). <em>Google&#8217;s Gemini 3.5 Flash Cyber becomes a vulnerability hunter</em>. https://www.helpnetsecurity.com/2026/07/22/google-gemini-3-5-flash-cyber-model/</p><p>Kahn, J. (2026, July 21). <em>OpenAI says its AI models escaped from a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation</em>. Fortune. https://fortune.com/2026/07/21/openai-says-ai-models-escaped-control-hacked-hugging-face/</p><p>MarkTechPost. (2026, July 22). <em>Anthropic releases Claude Security plugin for Claude Code in beta: A multi-agent vulnerability scanner that runs in your terminal</em>. https://www.marktechpost.com/2026/07/22/anthropic-releases-claude-security-plugin-for-claude-code-in-beta-a-multi-agent-vulnerability-scanner-that-runs-in-your-terminal/</p><p>Narayan, O., Singh, R., &amp; Baskar, P. (2026, July). <em>The Chronos vulnerability: A taxonomy of temporal persistence and memory-based deception in agentic AI</em> (arXiv:2607.19433). arXiv. https://arxiv.org/abs/2607.19433</p><p>SecurityOnline. (2026, July 20). <em>Linux kernel publishes 440 CVE advisories in 24 hours as AI accelerates bug hunting</em>. https://securityonline.info/linux-kernel-cves/</p><p>SiliconANGLE. (2026, July 20). <em>Neo Security bags $100M to build the secure control layer for enterprise AI agents</em>. https://siliconangle.com/2026/07/20/neo-security-bags-100m-build-secure-control-layer-enterprise-ai-agents/</p><p>The Hacker News. (2026, July 22). <em>Google launches Gemini 3.5 Flash Cyber AI to find and fix software vulnerabilities</em>. https://thehackernews.com/2026/07/google-launches-gemini-35-flash-cyber.html</p><p>Willison, S. (2026, July 22). <em>OpenAI&#8217;s accidental cyberattack against Hugging Face is science fiction that happened</em>. https://simonwillison.net/2026/Jul/22/openai-cyberattack/</p><p>XenoSpectrum. (2026, July). <em>Tracking the reality behind 440 Linux kernel CVE disclosures: Bulk assignment to already-fixed bugs and AI auditing</em>. https://xenospectrum.com/en/linux-kernel-cve-batch-ai-review/</p>]]></content:encoded></item><item><title><![CDATA[AI Defense Matrix: 194 Products Protect. Two Recover.]]></title><description><![CDATA[The AI Defense Matrix catalog maps 239 AI security products. Nineteen of the 48 cells sit empty or near-empty, every one in Govern, Respond, or Recover.]]></description><link>https://www.rockcybermusings.com/p/ai-defense-matrix-coverage-gap</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/ai-defense-matrix-coverage-gap</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 21 Jul 2026 12:51:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oaJH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oaJH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oaJH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!oaJH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!oaJH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!oaJH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oaJH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4069471,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/207800487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oaJH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!oaJH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!oaJH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!oaJH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a01dbf7-2ca5-4891-a8c3-2c3626fbe511_2048x2048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Open the AI Defense Matrix catalog and filter to Recover. You get two products out of 239. Filter to Protect and you get 194. Same catalog, same day, same 239 products. I pulled the raw data file yesterday and recounted every cell myself, because a gap that wide usually means somebody miscounted. Nobody miscounted.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/ai-defense-matrix-coverage-gap?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/ai-defense-matrix-coverage-gap?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2><span>What The AI Defense Matrix Measures</span></h2><p><span>Lenny Zeltser and Sounile Yu built a grid with eight AI-specific asset classes down the side, the six NIST CSF 2.0 functions across the top, 48 cells where they meet. </span><a href="https://aidefensematrix.com"><span>The AI Defense Matrix</span></a><span> launched May 11, 2026 under CC BY-SA 4.0, as the security-for-AI companion to Yu&#8217;s </span><a href="https://cyberdefensematrix.com"><span>Cyber Defense Matrix</span></a><span>.</span></p><p><span>The rule they used to decide what gets a row is what makes the thing useful. A row exists only if defending that asset requires AI-specific considerations, processes, or tools. If a generic security approach handles the defense well, it belongs in the Cyber Defense Matrix instead. Raw GPUs and containers got cut for that exact reason, and the change log says so out loud, which is more than most frameworks bother to do.</span></p><p><span>That discipline is why counting the cells means anything. </span><a href="https://catalog.aidefensematrix.com"><span>The companion catalog</span></a><span> places 239 products across those 48 cells, sourced from public information, dated, open to community correction. It catalogs AI-specific security products. Your SOAR platform and your backup vendor aren&#8217;t in it, and they shouldn&#8217;t be. What it measures is whether the industry has built an AI-specific answer for each cell.</span></p><p><span>For most of the grid, it has. For one corner of it, nothing showed up at all.</span></p><h2><span>Run The Count Yourself</span></h2><p><span>Nineteen of the 48 cells hold zero products or exactly one. All 19 sit in Govern, Respond, or Recover.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iETC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iETC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png 424w, https://substackcdn.com/image/fetch/$s_!iETC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png 848w, https://substackcdn.com/image/fetch/$s_!iETC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png 1272w, https://substackcdn.com/image/fetch/$s_!iETC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iETC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png" width="1456" height="918" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:918,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:258258,&quot;alt&quot;:&quot;Coverage grid of the AI Defense Matrix showing eight AI asset classes against six NIST CSF 2.0 functions, with the 19 cells holding zero or one product outlined in coral and all of them falling in the Govern, Respond, and Recover columns.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/207800487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Coverage grid of the AI Defense Matrix showing eight AI asset classes against six NIST CSF 2.0 functions, with the 19 cells holding zero or one product outlined in coral and all of them falling in the Govern, Respond, and Recover columns." title="Coverage grid of the AI Defense Matrix showing eight AI asset classes against six NIST CSF 2.0 functions, with the 19 cells holding zero or one product outlined in coral and all of them falling in the Govern, Respond, and Recover columns." srcset="https://substackcdn.com/image/fetch/$s_!iETC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png 424w, https://substackcdn.com/image/fetch/$s_!iETC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png 848w, https://substackcdn.com/image/fetch/$s_!iETC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png 1272w, https://substackcdn.com/image/fetch/$s_!iETC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc83d2a3-5945-412d-826e-87f027829d95_3900x2460.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: 239 AI Security Products On 48</figcaption></figure></div><p><span>Run it the other way and the picture gets starker. Every single cell in Identify, Protect, and Detect holds at least three products. The floor is three for Identify, seven for Protect, six for Detect. Not one thin cell anywhere in those three columns. Then you cross into Govern, and the floor drops to zero. Same for Respond. Same for Recover.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LX3B!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LX3B!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png 424w, https://substackcdn.com/image/fetch/$s_!LX3B!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png 848w, https://substackcdn.com/image/fetch/$s_!LX3B!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png 1272w, https://substackcdn.com/image/fetch/$s_!LX3B!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LX3B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png" width="1456" height="777" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:777,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:153946,&quot;alt&quot;:&quot;Bar chart counting cells that hold zero products or one in each NIST CSF function, showing zero such cells in Identify, Protect, and Detect, against five in Govern, six in Respond, and all eight in Recover.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/207800487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart counting cells that hold zero products or one in each NIST CSF function, showing zero such cells in Identify, Protect, and Detect, against five in Govern, six in Respond, and all eight in Recover." title="Bar chart counting cells that hold zero products or one in each NIST CSF function, showing zero such cells in Identify, Protect, and Detect, against five in Govern, six in Respond, and all eight in Recover." srcset="https://substackcdn.com/image/fetch/$s_!LX3B!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png 424w, https://substackcdn.com/image/fetch/$s_!LX3B!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png 848w, https://substackcdn.com/image/fetch/$s_!LX3B!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png 1272w, https://substackcdn.com/image/fetch/$s_!LX3B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa1acead-103a-4314-b0bc-df20a8956b4e_3600x1920.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: How Many Cells Per Column Are Empty Or Nearly Empty</figcaption></figure></div><p><span>Before anyone writes to tell me thin rows explain thin cells, they don&#8217;t. Runtime AI Data is the fattest row on the grid, with 244 product placements across it, and it still shows blanks in Govern and Recover. AI Model carries 116 placements and covers all six functions. Row size doesn&#8217;t predict where the holes are. Column does.</span></p><p><span>One number needs a footnote so nobody catches me playing games with it. When I say 194 products do Protect, that&#8217;s how many products carry a Protect tag on at least one asset class. Those same products land in 312 asset-specific Protect cells, because a product covering three rows counts once in the first number and three times in the second. Both numbers are real. They answer different questions.</span></p><p><span>Now, I&#8217;d be remiss if I didn&#8217;t call out that my employer sits in this catalog with five cells: agent identity discovery, plus protection and detection on runtime AI data and orchestration tools. All five land in the middle three columns. I&#8217;m not going to pretend that&#8217;s a coincidence or a shortcoming particular to us. It&#8217;s what nearly every product in the AI security market looks like when you place it on this grid, including the ones you already bought, and it&#8217;s the shape of the problem rather than the shape of any one vendor&#8217;s roadmap.</span></p><h2><span>Two Products Cover Recover</span></h2><p><span>Here they are, by name. Hirundo does machine unlearning, stripping memorized data and jailbreak behavior out of a trained model without a full retrain, and the catalog places it in Recover for AI Model. Rubrik Agent Cloud rewinds destructive agent actions, and it holds Recover for AI Agent Identities and for AI-Generated Code.</span></p><p><span>That&#8217;s the column. Two products, three cells, and no cell in Recover holds more than one product.</span></p><p><span>Respond does better and still doesn&#8217;t do well. Nine products across five cells. Four of the nine hold their Respond coverage on the AI Agent Identities row, and all four are non-human identity platforms: Astrix, Clutch, Permiso, and Vorlon. Credential revocation and agent quarantine are the one response capability the market ships in any volume at all.</span></p><p><span>Read the actual cell text and the two columns split cleanly. Respond is written in containment verbs: credential revocation, agent quarantine, session termination, RAG source isolation, plugin disable, shadow AI takedown. Recover is written in undo verbs: model version restore, dataset restore from golden copies, vector DB restore, re-indexing, prompt rollback, agent identity re-provisioning.</span></p><p><span>Then there&#8217;s the detail I keep coming back to. Three cells on the entire grid are labeled generic by the authors themselves. Generic container IR. Generic platform restore. Generic network failover. All three sit in Respond and Recover. Nowhere else on 48 cells do Zeltser and Yu tell you the traditional answer will do. That admission is the most honest thing on the page, and it&#8217;s also a map of where the AI-specific work hasn&#8217;t happened yet.</span></p><p><span>Undo assumes a lot. An agent that ran for six hours holding delegated credentials across a dozen tools has sent emails, opened pull requests, written to a CRM, posted in channels, and fired webhooks that triggered things you don&#8217;t have visibility into. Restore the platform, and none of that comes back. Those effects live in separate systems with separate owners and no shared transaction boundary, which means there&#8217;s nothing to roll back to. I&#8217;ve written before about why </span><a href="https://www.rockcybermusings.com/p/i-agent-authentication-authorization-gap"><span>least agency beats least privilege</span></a><span> once agents can act, and this column is that argument showing up as somebody else&#8217;s control taxonomy.</span></p><h2><span>Why The Market Sells The Middle</span></h2><p><span>I had the wrong explanation for this at first. My first read was that the empty columns are the architectural ones, and the market only sells what bolts on. Identify killed that theory in about a minute. Discovering AI agents across an enterprise is an architecture problem too, and the market sells 146 products against it.</span></p><p><span>That&#8217;s the difference between watching a system and reaching into it.</span></p><p><span>Identify and Detect sell because a product can do the work from outside. Agentless posture management crawls your cloud and finds unregistered models, and you change nothing to let it. Protect sells at 194 for a reason that surprised me until I read the product descriptions: what the market calls Protect for AI is overwhelmingly inline filtering at a chokepoint. Gateways, proxies, browser extensions, guardrails on prompts and responses. That&#8217;s still exterior work. A filter inspects something crossing a boundary and decides yes or no.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pf1f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pf1f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png 424w, https://substackcdn.com/image/fetch/$s_!pf1f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png 848w, https://substackcdn.com/image/fetch/$s_!pf1f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png 1272w, https://substackcdn.com/image/fetch/$s_!pf1f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pf1f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png" width="3900" height="2100" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2100,&quot;width&quot;:3900,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:459873,&quot;alt&quot;:&quot;Two-zone diagram separating the AI security functions a product delivers from outside a system, Identify at 146 products, Protect at 194 and Detect at 191, from the functions that require reaching into it, Govern at 11, Respond at nine and Recover at two.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/207800487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67312b30-7e4e-4331-a4f8-d814b2559c1c_3900x2100.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two-zone diagram separating the AI security functions a product delivers from outside a system, Identify at 146 products, Protect at 194 and Detect at 191, from the functions that require reaching into it, Govern at 11, Respond at nine and Recover at two." title="Two-zone diagram separating the AI security functions a product delivers from outside a system, Identify at 146 products, Protect at 194 and Detect at 191, from the functions that require reaching into it, Govern at 11, Respond at nine and Recover at two." srcset="https://substackcdn.com/image/fetch/$s_!pf1f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png 424w, https://substackcdn.com/image/fetch/$s_!pf1f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png 848w, https://substackcdn.com/image/fetch/$s_!pf1f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png 1272w, https://substackcdn.com/image/fetch/$s_!pf1f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce355f26-1ce2-424d-88d9-da5d84c51a02_3900x2100.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: What The Market Sells And What You Build</figcaption></figure></div><p><span>Every containment verb in the Respond column carries a precondition somebody had to build. You terminate a session where sessions exist and can be terminated. You revoke a credential where it was scoped and revocable in the first place. You quarantine an agent where the agent runs somewhere quarantinable. A vendor can sell you the button. Nobody can sell you the wiring behind it.</span></p><p><span>One row breaks that pattern, and it&#8217;s the exception worth studying rather than the one that sinks the argument. AI Agent Identities carries 56 products under Protect, and the cell text reads agent OAuth, capability scoping, and short-lived credentials. That&#8217;s architecture sold as a product, which is the thing I said the market doesn&#8217;t do.</span></p><p><span>Look at what happened behind it. AI Agent Identities and AI Model are the only two rows on the whole grid with coverage in all six functions, Respond and Recover included. Every other row has at least one blank. Both of those rows got a market around the underlying primitive first, non-human identity tooling in one case and model registries and provenance in the other. Once the asset became addressable, once agents had scoped identities and models had versions and inventories, the containment columns filled in behind them. Four products doing agent quarantine exist because 56 products turned agents into principals you can scope.</span></p><p><span>That sequence is the roadmap, and it took me a while to see it. Response capability doesn&#8217;t show up because vendors decide to care about incidents. It shows up after somebody makes the asset addressable, and containment becomes something you can build on top of that. Runtime AI Data carries 244 placements and zero Recover coverage, because nobody has made a prompt, a RAG chunk, or an agent&#8217;s persistent memory into a thing with a version and an owner you can roll back to. Do that work and the column can fill. Skip it and vendor attention won&#8217;t rescue you.</span></p><p><span>Govern is the cleanest proof, and it settles a question I went back and forth on. Read all eight Govern cells and every one is a standard, a policy, or an evaluation decision. AI platform standards. AI application governance. AI coding standards and code-review policy. AI egress policy. Model selection and provider evaluation. Dataset provenance and licensing policy. Prompt and RAG policy. AI agent identity policy and authorization standards.</span></p><p><span>None of that is a product category. Zeltser and Yu tell you as much in their own instructions, where the gap analysis asks whether process, technology, or both cover each cell, and then tells you to start with Govern. Eleven products carry a Govern tag, and five of the eight rows have zero. Credo AI, OneTrust, Cranium, Holistic AI, and LatticeFlow are all in there and correctly tagged. They help you document and evidence a policy. Writing the policy stays yours.</span></p><h2><span>The Prediction The AI Defense Matrix Sets Up</span></h2><p><span>Here&#8217;s where I part company with the authors, and I want to be precise about it because they are probably right.</span></p><p><span>Their guidance to vendors says to treat thinly covered cells as opportunities for differentiation and new products that solve underserved needs. That&#8217;s a forecast because the thin columns fill because the market notices the gap and builds into it. My read is that Recover and most of Respond stay thin, because the missing capability isn&#8217;t a product somebody forgot to build. It&#8217;s a property of systems that were designed without it.</span></p><p><span>Put a number on it. I&#8217;d say 70% that the Recover column holds fewer than ten products in July 2028, and under 20% that it ever resembles what Protect looks like today. I&#8217;m at 70 rather than 90 because I&#8217;ve been wrong about market timing before. Nobody sold cloud posture management in 2016 either, and a July 2026 snapshot of a young category can look identical to a permanent boundary.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lTFY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lTFY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png 424w, https://substackcdn.com/image/fetch/$s_!lTFY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png 848w, https://substackcdn.com/image/fetch/$s_!lTFY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png 1272w, https://substackcdn.com/image/fetch/$s_!lTFY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lTFY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png" width="1456" height="801" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:801,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:219116,&quot;alt&quot;:&quot;Line chart contrasting two paths for the number of products covering Recover between July 2026 and July 2028, a flat structural path staying under ten and a market-fills path climbing past eighty, with the falsifying condition named.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/207800487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart contrasting two paths for the number of products covering Recover between July 2026 and July 2028, a flat structural path staying under ten and a market-fills path climbing past eighty, with the falsifying condition named." title="Line chart contrasting two paths for the number of products covering Recover between July 2026 and July 2028, a flat structural path staying under ten and a market-fills path climbing past eighty, with the falsifying condition named." srcset="https://substackcdn.com/image/fetch/$s_!lTFY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png 424w, https://substackcdn.com/image/fetch/$s_!lTFY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png 848w, https://substackcdn.com/image/fetch/$s_!lTFY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png 1272w, https://substackcdn.com/image/fetch/$s_!lTFY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F384e10fb-108c-43f6-9906-ad67b62d4355_3600x1980.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: A Claim You Can Check In Two Years</figcaption></figure></div><p><span>The falsifier is specific, so hold me to it. Watch Rubrik. If three more vendors ship agent-action reversal and it works without you rerouting your agents through their control plane first, I&#8217;m wrong and the market solved something I said it structurally couldn&#8217;t. If the only way it ever works is routing your agents through somebody&#8217;s plane, that&#8217;s the architectural decision doing the work, and the product is the part you bolt on afterward.</span></p><p><span>Meanwhile, NIST IR 8596, the Cyber AI Profile everyone&#8217;s waiting on, has been sitting as an initial preliminary draft since December 16, 2025, with the comment period closed since January 30, 2026. Seven months, no initial public draft. You need a working instrument now, and this grid is the best one anybody&#8217;s published.</span></p><p><span>Two products cover Recover for all of AI. One strips bad behavior out of model weights. The other rewinds agent actions it was routed through ahead of time. That&#8217;s the entire Recover column of the AI Defense Matrix, on a grid of 48 cells, in a market that has shipped 194 ways to protect you and almost nothing to put you back together.</span></p><h3><span>What To Do Next</span></h3><p><span>Download the matrix as a CSV, YAML, or Markdown and populate it for your own environment this week. Mark every cell covered, partial, or absent. Budget 90 minutes with whoever owns your agent platforms.</span></p><p><span>Then cover Identify, Protect, and Detect with your hand and look at what&#8217;s left. For each blank in Govern, Respond, or Recover, write down two things: the design decision that would have to change for that cell to fill, and the name of the person who can make that decision. Not the product that would fill it. The decision and the owner. That&#8217;s a build backlog and an ownership map, which is Create and Adapt work in the </span><a href="https://www.rockcyber.com/ai-strategy-and-governance"><span>CARE model</span></a><span>, and it&#8217;s a different conversation than a purchase order.</span></p><p><span>If you want the deeper argument on why authorization scope is a runtime decision rather than a connection-time checkbox, I wrote that up in </span><a href="https://www.rockcybermusings.com/p/mcp-authorization-scope-spec-gap"><span>the MCP authorization gap piece</span></a><span>, and </span><a href="https://www.rockcybermusings.com/p/autonomous-ai-agent-ransomware-jadepuffer"><span>JadePuffer</span></a><span> is what the Respond column being empty looks like when it happens to somebody. More at </span><a href="https://rockcybermusings.com"><span>rockcybermusings.com</span></a><span>, and the advisory work lives at </span><a href="https://www.rockcyber.com"><span>rockcyber.com</span></a><span>.</span></p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation continues at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span>.</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p>&#128073; As a bonus, check out my latest appearance on the <strong><a href="https://dtsr.buzzsprout.com/">Down The Security Rabbit Hole</a></strong> podcast, where we discuss the current state of AI in the&nbsp;SOC for cybersecurit<a href="https://www.youtube.com/hashtag/cybersecurity"><span>y</span></a>&nbsp;and specifically how it differentiates from the promise of SOAR.</p><div id="youtube2-5YIpZuQLTMg" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;5YIpZuQLTMg&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/5YIpZuQLTMg?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><p><br></p>]]></content:encoded></item><item><title><![CDATA[Weekly Musings Top 10 AI Security Wrapup: Issue 46 July 10 -July 16, 2026]]></title><description><![CDATA[The week agentic AI security got real: an AI agent breached Hugging Face, a red team cracked GPT-5.6 in six hours, and China's first AI-agent rules went live.]]></description><link>https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260710-20260716</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260710-20260716</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Fri, 17 Jul 2026 12:51:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3uJE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3uJE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3uJE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!3uJE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!3uJE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!3uJE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3uJE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1233556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/207367589?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3uJE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!3uJE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!3uJE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!3uJE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68f8f3a8-7fd2-4881-a261-da3d60b137a8_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We&#8217;ve warned boards that autonomous AI attacks were coming. We were proven right once again. An AI agent broke into Hugging Face and ran the intrusion itself. A government red team cracked a new frontier model&#8217;s cyber safeguards in six hours. The forecast became the incident report.</p><p>The through-line rattles anyone still treating AI as a feature bolted onto the business. Attackers now point capable models at your environment and let them plan, adapt, and act at machine speed. China switched on the first real rulebook for AI agents. Two labs dropped open-weight frontier models anyone can download. Microsoft shipped its largest patch load ever and credited its own AI. I cut the stories that predate July 10. Here is what landed.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260710-20260716?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260710-20260716?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>1. Hugging Face Discloses a Breach Run End to End by an AI Agent</h3><p>On July 16, 2026, Hugging Face disclosed that an autonomous AI agent breached part of its production infrastructure days earlier (Hugging Face, 2026). A malicious dataset abused two code-execution paths, escalated to node-level access, and moved laterally over a weekend. The company found access to internal datasets and service credentials, with no public models touched. It reconstructed the timeline with its own LLM analysis over 17,000 attacker events (NHIMG, 2026).</p><p><strong>Why it matters</strong></p><ul><li><p>The first high-profile platform breach run end to end by an agent.</p></li><li><p>The foothold was a poisoned dataset, not a phished user.</p></li><li><p>Safety guardrails blocked the defenders, not the attacker.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Sandbox dataset-processing workers and scope their credentials.</p></li><li><p>Pre-stage an open-weight model that you control for incident response.</p></li><li><p>Rehearse weekend and off-hours detection.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I have sat through a hundred tabletops where someone called this scenario fiction. Hand them the disclosure. The attacker chained known mistakes into one operation, and only the driver was new. Guardrails blocked the responders while the attacker&#8217;s stripped-down model faced none. Stage your break-glass model, rotate the secrets your workers touch, and expect the next probe on a Saturday. I have argued this blast-radius math at rockcybermusings.com for a while.</p><h3>2. UK Government Red Team Cracks GPT-5.6 Cyber Safeguards in Six Hours</h3><p>On July 10, 2026, the UK AI Security Institute reported universal jailbreaks in OpenAI&#8217;s freshly released GPT-5.6 (Fortune, 2026). Testers defeated the controls meant to stop the model from hacking, then used it to break into systems on its own. AISI built the jailbreak in roughly six hours, and it held across OpenAI&#8217;s malicious-query set (Technobezz, 2026). OpenAI pointed to its layered safeguards and a rapid remediation process.</p><p><strong>Why it matters</strong></p><ul><li><p>A government agency beat a frontier lab&#8217;s cyber guardrails in an afternoon.</p></li><li><p>These jailbreaks enable agentic offense rather than one-off bad answers.</p></li><li><p>Pre-release government testing caught it before public exposure.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Wrap your own boundaries around any model with tool access.</p></li><li><p>Make vendor red-team results a written procurement gate.</p></li><li><p>Assume any model exposed to untrusted input can turn offensive.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Six hours. That is the number I want every executive to sit with. A red team needed less than a workday to turn a flagship model into an exploit-writing partner. I am a Bayesian, and my prior that guardrails hold under a determined attacker dropped again. Vendor safety layers are speed bumps, worthless against the motivated. The redeeming part is a government caught this before release. Put the vendor&#8217;s evaluation posture in your contract.</p><h3>3. China&#8217;s Rulebook for AI Agents Takes Effect, the First of Its Kind</h3><p>On July 15, 2026, China&#8217;s Implementation Opinions on intelligent agents became enforceable, the world&#8217;s first dedicated regulatory category for AI agents (IAPP, 2026). The rules define agents as systems capable of autonomous perception, memory, decision-making, and execution. They establish a three-tier decision-authorization structure and mandate filing, testing, and product-recall provisions in sensitive sectors (Global Law Experts, 2026). The same date pushed platforms like Doubao and Qwen to shut down companion features.</p><p><strong>Why it matters</strong></p><ul><li><p>The first binding definition of an AI agent came out of Beijing.</p></li><li><p>China operations with agent deployments face immediate filing duties.</p></li><li><p>Authorization tiers force you to declare what an agent decides alone.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>In China, map every agent to the authorization and filing rules now.</p></li><li><p>Build an agent inventory with decision rights and override thresholds.</p></li><li><p>Align your internal agent taxonomy to the definition.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I do not love handing Beijing the pen, but credit where it is due. They defined an AI agent while the rest of us argued about whether agents are high-risk by default. Definitions are power, and the clearest legal one now lives in a Chinese regulation. The practical bite is the tiers: what may this agent decide without a human in the loop? I ask that in every board session I run through rockcyber.com and get blank stares more often than answers. Write down your agents&#8217; decision rights now.</p><h3>4. xAI&#8217;s Grok Build CLI Shipped Entire Repositories and Secrets to the Cloud</h3><p>Around July 13 and 14, 2026, a researcher showed that version 0.2.93 of xAI&#8217;s Grok Build coding CLI uploaded entire Git repositories to a Google Cloud Storage bucket rather than the files a task needed (Tech Times, 2026). For a 12 GB repository, the tool pushed more than 5 GB while model traffic came to 192 KB. The upload included the full commit history, so deleted files went too. Files holding credentials uploaded without redaction, and opting out did not stop the sweep (Crypto Briefing, 2026).</p><p><strong>Why it matters</strong></p><ul><li><p>Your code and secrets can leave as a side effect of a coding assistant working.</p></li><li><p>The opt-out control did nothing. Privacy toggles are marketing until proven.</p></li><li><p>Full commit-history upload means deleted secrets are back in play.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Block or quarantine Grok Build CLI until you verify what it transmits.</p></li><li><p>Gate every agentic coding tool on evidence of what it sends.</p></li><li><p>Rotate exposed secrets and move credentials into a manager.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This one made me put my coffee down, because the privacy toggle was theater. Someone shipped a product where do-not-send-my-data was a suggestion the software ignored. We only know because a researcher watched the packets. How many other agentic tools in your environment do the same and nobody has checked? Agentic developer tools are the new shadow IT, with your whole codebase as the blast radius. Gate approval on proof of behavior, and get secrets out of source control.</p><h3>5. Unpatched Claude for Chrome Flaw Lets Rogue Extensions Read Your Gmail</h3><p>Mid-week, researchers at Manifold Security detailed two flaws in Anthropic&#8217;s Claude for Chrome extension that let a malicious browser extension read a victim&#8217;s Gmail, Docs, and Calendar with about six lines of JavaScript (The Hacker News, 2026). The root cause is a handler that never checks whether a click came from a real user. Any extension with script access on claude.ai forges a click and fires a hardcoded prompt silently. Manifold reported the flaws in May 2026, yet the vulnerable code still shipped in version 1.0.80 on July 7 (CSO Online, 2026).</p><p><strong>Why it matters</strong></p><ul><li><p>A browser AI agent with mailbox access becomes an exfiltration tool.</p></li><li><p>The flaw maps to indirect prompt injection and excessive agency.</p></li><li><p>Reported in May, still exploitable in July. Resolved did not mean fixed.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory AI browser extensions with mail access, and disable act-without-asking mode.</p></li><li><p>Treat browser AI agents as privileged identities.</p></li><li><p>Demand proof a fix reached production, not that a ticket closed.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>The detail that stuck with me is the gap between a ticket marked resolved and code that never moved. Someone closed it, everyone relaxed, and eight releases later the hole is still open. I have lived that failure mode in every large program I have run. The browser holds your sensitive data already decrypted and logged in, and you handed an agent permission to read your mail. The guardrail against abuse was a check the developers forgot to write. Treat every browser agent as a privileged service account, and ask to see the diff.</p><h3>6. Microsoft Patches a Record 570 Flaws as AI Reshapes Both Sides of the Fight</h3><p>On July 14, 2026, Microsoft shipped the largest Patch Tuesday in its history, fixing a record 570 vulnerabilities, including three zero-days with two exploited in the wild (BleepingComputer, 2026). Microsoft tied the surge to its own AI, crediting a vulnerability scanner it calls MDASH that finds flaws in Windows components before attackers reach them. The number runs roughly four times the same month a year earlier (Krebs on Security, 2026). The exploited zero-days hit Active Directory Federation Services and SharePoint Server (TechCrunch, 2026).</p><p><strong>Why it matters</strong></p><ul><li><p>AI now finds vulnerabilities at a scale that breaks your patch cadence.</p></li><li><p>The tools that help Microsoft find bugs help attackers weaponize them in hours.</p></li><li><p>Two exploited zero-days sit in identity infrastructure.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Prioritize the exploited AD FS and SharePoint zero-days now.</p></li><li><p>Shorten patch deferral windows for internet-facing and identity systems.</p></li><li><p>Invest in runtime detection, not only patching.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Five hundred and seventy. No human triage process was built for a number like that, and Microsoft is honest that AI drove the spike. Their agent found the bugs before the bad guys, who aim the same tools at the same code. Offense and defense now both scale with compute, and the loser moves at human speed. The energy and manufacturing clients I worry about cannot patch on this timeline, so they lean on segmentation and detection tuned for exploitation behavior. Plan for the flood, because it is the baseline now.</p><h3>7. Thinking Machines Ships Inkling, a 975-Billion-Parameter Open-Weight Model That Fine-Tunes Itself</h3><p>On July 15, 2026, Mira Murati&#8217;s Thinking Machines Lab released Inkling, an open-weights mixture-of-experts transformer with 975 billion total parameters and 41 billion active (TechCrunch, 2026). It carries a context window up to 1 million tokens. The headline capability was a demonstration in which Inkling fine-tuned itself, autonomously writing, running, and evaluating its own fine-tuning job (Axios, 2026). Full weights mean anyone can download, inspect, and customize the model.</p><p><strong>Why it matters</strong></p><ul><li><p>A frontier-scale model with open weights lands in anyone&#8217;s hands.</p></li><li><p>Self-directed fine-tuning collapses the effort to specialize a model.</p></li><li><p>Open weights on your own hardware is what defenders wanted this week.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Add open-weight models to your threat model as unrestricted attacker capability.</p></li><li><p>Treat internal self-fine-tuning as a change-management event.</p></li><li><p>Give defenders a sanctioned open-weight option on controlled infrastructure.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Open weights are a genuine double-edged sword, and I refuse to pretend otherwise for either camp. The same day Hugging Face said defenders need a self-hosted model, Thinking Machines handed everyone a frontier-scale one. That helps blue teams and it helps the people building offense models with the safety sanded off. My probability that abliterated versions show up in attack tooling within months is high. I am not in the ban-open-models camp. What I want is honesty about the trade.</p><h3>8. Kimi K3 Pushes an Open-Weight Model to the Frontier, From China</h3><p>On July 16, 2026, Moonshot AI launched Kimi K3, a roughly 2.8-trillion-parameter mixture-of-experts model. Its 1-million-token context window targets long-horizon coding and agent workloads (TechCrunch, 2026). In blind testing by the evaluation platform Arena, developers preferred Kimi over every leading US model for front-end coding, including Anthropic&#8217;s Fable 5 and OpenAI&#8217;s GPT-5.6 Sol (Axios, 2026). Moonshot committed to releasing the open weights on July 27.</p><p><strong>Why it matters</strong></p><ul><li><p>A Chinese lab reached rough parity and chose to open the weights.</p></li><li><p>Frontier-level coding in open weights lowers the floor for autonomous tooling.</p></li><li><p>Enterprises will feel pressure to adopt the cheapest capable model.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Decide your policy on Chinese-origin open-weight models now.</p></li><li><p>Factor the July 27 weight release into your threat model.</p></li><li><p>Evaluate data residency before routing sensitive code through a hosted Kimi endpoint.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Two open-weight frontier models landed in one week. One came from a former OpenAI CTO, one from Beijing, both matching or beating the closed US flagships. The story that closed labs hold an unbridgeable lead died quietly this week. When developers in a blind test prefer Kimi over Fable 5 and GPT-5.6, price and openness win the next procurement cycle. The question I keep dragging boards toward: what is your policy on where your models come from, and did you write it yet? Decide deliberately, because the July 27 weight drop turns this into an anyone-can-run-it problem.</p><h3>9. Meta Pulls an Instagram AI Feature That Copied People&#8217;s Likenesses Without Clear Consent</h3><p>After a short, loud backlash, Meta removed the Muse Image feature that let users generate AI images by referencing public Instagram accounts, with the reversal landing around July 13 and 14, 2026 (TechRepublic, 2026). Launched July 7, it used public accounts&#8217; posted photos as references. Adult public accounts were swept in by default unless a user disabled a reuse setting, and Meta&#8217;s own policy said people would not be notified when someone used their likeness (Variety, 2026). Meta said the feature missed the mark and pulled it.</p><p><strong>Why it matters</strong></p><ul><li><p>A trillion-dollar platform shipped opt-out-by-default likeness reuse and retreated in a week.</p></li><li><p>Likeness reuse is what the EU AI Act transparency rules move to govern.</p></li><li><p>The reputational half-life of a bad AI consent decision is now days.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Default any AI feature using third-party likenesses to opt-in with notice.</p></li><li><p>Map likeness features against the EU AI Act obligations arriving August 2.</p></li><li><p>Put a consent review gate in front of AI feature launches.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Meta is not a naive startup that stumbled into a consent problem. They have armies of lawyers and privacy engineers, and they still shipped default-on likeness reuse with no notice to the people whose faces got borrowed. The reflex inside big platforms remains take first, apologize later. What changed is the clock, because the backlash cost them the feature in a week. I tell boards that consent is not a checkbox at the bottom of a settings page nobody opens, and this is the case study. Design for affirmative consent now.</p><h3>10. Check Point&#8217;s AI Security Report Puts Numbers to the Machine-Speed Threat</h3><p>On July 15, 2026, Check Point published its AI Security Report 2026, documenting a year in which attackers ran exploitation workflows autonomously, generating thousands of commands with minimal human direction (Help Net Security, 2026). Adversaries acquire AI capability by abusing commercial models, stealing credentials, or self-hosting open-source models, then strip the safety controls. The report flags jailbreaks planted in agent configuration files like CLAUDE.md that load every session and stay active until removed. On identity, trained observers spotted AI-generated faces only about 41% of the time (Check Point, 2026).</p><p><strong>Why it matters</strong></p><ul><li><p>The autonomous-attacker pattern is now measured across a year of real incidents.</p></li><li><p>Jailbreaks in config files keep an agent compromised across sessions.</p></li><li><p>Trained analysts spot fake faces 41% of the time, so identity checks fail.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Add integrity monitoring for agent configuration files.</p></li><li><p>Move identity verification beyond face and voice for high-value actions.</p></li><li><p>Measure your own high-risk AI interaction rate.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I am wary of vendor reports, so I read this one for the load-bearing data. The number that survived my skepticism is the 41%. When trained people flip a coin and do worse at spotting a fake face, every identity control built on I-can-tell-it-is-you is obsolete. I have told financial and energy clients for two years that voice and video prove nothing anymore. The config-file jailbreak deserves more attention than it will get, because a poisoned CLAUDE.md is patient persistence most stacks never inspect. Instrument the files your agents read at startup.</p><h3>The One Thing You Won&#8217;t Hear About But You Need To</h3><h3>11. Defenders Turn Prompt Injection Into a Weapon: Tracebit&#8217;s &#8220;Context Bombs&#8221;</h3><p>On July 14, 2026, researchers at Tracebit published a defensive twist that flips prompt injection against the attacker (Help Net Security, 2026). Rather than hijacking an AI, they plant short strings called context bombs inside decoy resources, crafted to trip the safety guardrails of an offensive AI agent and stop it mid-intrusion. They tested five leading models across 152 trials in an imitation AWS environment. Opus 4.8 reached full account admin in 93% of clean runs and failed every single time once a context bomb was in play (Tracebit, 2026).</p><p><strong>Why it matters</strong></p><ul><li><p>Defenders finally have a tool that turns the attacker&#8217;s safety training into a trap.</p></li><li><p>Context bombs disrupt the intrusion in progress and buy back reaction time.</p></li><li><p>The technique layers on top of canary and deception controls.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Pilot context bombs inside honeytokens and decoy secrets.</p></li><li><p>Pair them with canaries so a tripped bomb also raises an alert.</p></li><li><p>Track abliterated models, which may ignore the bomb.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This is the story that made me smile all week, and almost nobody outside the research crowd is talking about it. For two years the prompt-injection conversation has been a funeral dirge: the flaw cannot be fixed, abandon hope. Tracebit asked the contrarian question, why not turn the attacker&#8217;s guardrails into a landmine in your own environment? Opus 4.8 going from a 93% admin-access rate to zero resets a prior in a hurry. The caveat matters, because an attacker running an abliterated model may walk right past the bomb. I would still rather hand my defenders a working trap today than wait for a fix that is never coming.</p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation continues at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span>.</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, check our AMA on the </span><strong><a href="https://genai.owasp.org/resource/state-of-agentic-ai-security-and-governance/">2026 OWASP GenAI Security Project State of Agentic AI Security and Governance report</a></strong><span> with me and the other co-leads (it was live, so start at time marker 09:45)</span></p><div id="youtube2-jK1Z7Z6zlW0" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;jK1Z7Z6zlW0&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/jK1Z7Z6zlW0?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><em>The views and opinions expressed in RockCyber Musings are my own and do not represent the positions of my employer or any organization I&#8217;m affiliated with.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><h2>References</h2><p>Axios. (2026, July 15). <em>Mira Murati&#8217;s Thinking Machines debuts its first AI model</em>. Axios. https://www.axios.com/2026/07/15/mira-murati-thinking-machines-open-weight-model-inkling</p><p>Axios. (2026, July 16). <em>China&#8217;s open-weight Kimi model stuns AI world with frontier-level results</em>. Axios. https://www.axios.com/2026/07/16/moonshot-kimi-ai-china-model-openai-anthropic</p><p>BleepingComputer. (2026, July 14). <em>Microsoft July 2026 Patch Tuesday fixes massive 570 flaws, 3 zero-days</em>. BleepingComputer. https://www.bleepingcomputer.com/news/microsoft/microsoft-july-2026-patch-tuesday-fixes-massive-570-flaws-3-zero-days/</p><p>Check Point Software Technologies. (2026, July 15). <em>AI Security Report 2026</em>. Check Point. https://www.helpnetsecurity.com/2026/07/15/check-point-ai-security-report-2026/</p><p>Crypto Briefing. (2026, July 14). <em>xAI Grok Build CLI caught uploading private code and secrets to Google Cloud bucket</em>. Crypto Briefing. https://cryptobriefing.com/xai-grok-build-cli-private-code-leak/</p><p>CSO Online. (2026, July). <em>New bugs in Claude for Chrome allow extensions to abuse AI privileges</em>. CSO Online. https://www.csoonline.com/article/4197325/new-bugs-in-claude-for-chrome-allow-extensions-to-abuse-ai-privileges.html</p><p>Fortune. (2026, July 10). <em>U.K. agency finds &#8216;universal jailbreaks&#8217; unlock dangerous cyber capabilities of OpenAI&#8217;s GPT-5.6</em>. Fortune. https://fortune.com/2026/07/10/openai-gpt-5-6-sol-jailbreaks-cyber-attacks-similar-to-security-flaw-that-led-u-s-government-to-force-anthropic-to-disable-fable-5/</p><p>Global Law Experts. (2026, July). <em>A concise interpretation of the CAC&#8217;s Implementation Opinions on the standardised application and innovative development of AI agents</em>. Global Law Experts. https://globallawexperts.com/innovation-regulation-as-inseparable-objectives-a-concise-interpretation-of-the-cacs-implementation-opinions-on-the-standardised-application-innovative-development-of-ai-agents/</p><p>Help Net Security. (2026, July 14). <em>&#8220;Context bombs&#8221; can frustrate AI-driven attacks, researchers found</em>. Help Net Security. https://www.helpnetsecurity.com/2026/07/14/context-bombs-for-defensive-prompt-injection/</p><p>Help Net Security. (2026, July 15). <em>AI used to help plan the break-in, now it&#8217;s doing the break-in</em>. Help Net Security. https://www.helpnetsecurity.com/2026/07/15/check-point-ai-security-report-2026/</p><p>Hugging Face. (2026, July 16). <em>Security incident disclosure: July 2026</em>. Hugging Face. https://huggingface.co/blog/security-incident-july-2026</p><p>International Association of Privacy Professionals. (2026, July). <em>China&#8217;s new AI rules: Ethics, AI agents and anthropomorphic AI</em>. IAPP. https://iapp.org/news/a/china-s-new-ai-rules-ethics-ai-agents-and-anthropomorphic-ai</p><p>Krebs on Security. (2026, July). <em>Microsoft patches a record 570 security flaws</em>. Krebs on Security. https://krebsonsecurity.com/2026/07/microsoft-patches-a-record-570-security-flaws/</p><p>NHIMG. (2026, July). <em>Hugging Face breach</em>. Non-Human Identity Management Group. https://nhimg.org/hugging-face-breach</p><p>Technobezz. (2026, July). <em>U.K. AI Security Institute finds universal jailbreaks in OpenAI&#8217;s GPT-5.6 within hours</em>. Technobezz. https://www.technobezz.com/news/uk-ai-security-institute-finds-universal-jailbreaks-in-openais-gpt-56-within-hours</p><p>TechCrunch. (2026, July 15). <em>Thinking Machines amps up its bet against one-size-fits-all AI with its first open model, Inkling</em>. TechCrunch. https://techcrunch.com/2026/07/15/thinking-machines-amps-up-its-bet-against-one-size-fits-all-ai-with-its-first-open-model-inkling/</p><p>TechCrunch. (2026, July 15). <em>Microsoft patches record number of security vulnerabilities, citing its use of AI</em>. TechCrunch. https://techcrunch.com/2026/07/15/microsoft-patches-record-number-of-security-vulnerabilities-citing-its-use-of-ai/</p><p>TechCrunch. (2026, July 16). <em>Moonshot&#8217;s Kimi 3 is expected to close the gap with Anthropic&#8217;s Opus 4.8</em>. TechCrunch. https://techcrunch.com/2026/07/16/moonshots-upcoming-kimi-3-is-expected-to-close-the-gap-with-anthropics-opus-4-8/</p><p>TechRepublic. (2026, July). <em>Meta removes Muse Image Instagram feature after consent backlash</em>. TechRepublic. https://www.techrepublic.com/article/news-meta-scraps-muse-image-feature-after-launch/</p><p>Tech Times. (2026, July 14). <em>Grok Build shipped entire codebases to xAI cloud, privacy toggle did nothing</em>. Tech Times. https://www.techtimes.com/articles/320420/20260714/grok-build-shipped-entire-codebases-xai-cloud-privacy-toggle-did-nothing.htm</p><p>The Hacker News. (2026, July). <em>Researchers say Claude for Chrome flaw lets rogue extensions trigger Gmail reads</em>. The Hacker News. https://thehackernews.com/2026/07/claude-for-chrome-flaw-lets-other.html</p><p>Tracebit. (2026, July 14). <em>Context bombs: Stopping AI attackers in their tracks</em>. Tracebit Research. https://agentic.tracebit.com/context-bombs/</p><p>Variety. (2026, July). <em>Meta suspends Instagram AI image feature after days of backlash</em>. Variety. https://variety.com/2026/biz/news/meta-suspends-ai-image-instagram-feature-backlash-1236806989/</p>]]></content:encoded></item><item><title><![CDATA[Autonomous AI Agent Ransomware: JadePuffer’s Authorization Lesson]]></title><description><![CDATA[Autonomous AI agent ransomware arrived with JadePuffer. The real lesson is machine-identity authorization over exploit novelty, and the controls to build.]]></description><link>https://www.rockcybermusings.com/p/autonomous-ai-agent-ransomware-jadepuffer</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/autonomous-ai-agent-ransomware-jadepuffer</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Tue, 14 Jul 2026 12:50:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!fp9k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fp9k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fp9k!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png 424w, https://substackcdn.com/image/fetch/$s_!fp9k!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png 848w, https://substackcdn.com/image/fetch/$s_!fp9k!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png 1272w, https://substackcdn.com/image/fetch/$s_!fp9k!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fp9k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png" width="1456" height="1110" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1110,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3580211,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/206507719?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fp9k!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png 424w, https://substackcdn.com/image/fetch/$s_!fp9k!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png 848w, https://substackcdn.com/image/fetch/$s_!fp9k!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png 1272w, https://substackcdn.com/image/fetch/$s_!fp9k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ce1de92-6db0-471c-9243-d3d025d1a9c3_1975x1505.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Autonomous AI agent ransomware stopped being a research demo. </span><a href="https://www.sysdig.com/blog/jadepuffer-agentic-ransomware-for-automated-database-extortion"><span>Sysdig disclosed JadePuffer</span></a><span>, an operation in which a language model executed a full ransomware kill chain from recon to encryption, with no human at the keyboard. It used no zero-days. It used a year-old bug, default credentials, and machine identities you already own. Here&#8217;s what happened, and the one control that would have contained it.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/autonomous-ai-agent-ransomware-jadepuffer?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/autonomous-ai-agent-ransomware-jadepuffer?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2><strong><span>How The Agent Ran The Whole Chain</span></strong></h2><p><span>Langflow is a widely used open-source framework for building AI applications, and plenty of teams run it safely. This victim left a version with a year-old critical bug facing the open internet. That exposure is what the agent walked through. Once it had code execution on the host, it enumerated the machine and swept the environment for secrets in parallel: model provider keys, cloud credentials, crypto wallets, database credentials, and configuration files. It dumped Langflow&#8217;s own backing database, pulled the stored keys and user records, then deleted its staging files to clean up after itself.</span></p><p><span>Then it found the object store. Using the default credentials minioadmin:minioadmin, it listed every bucket, prioritized the ones named for infrastructure state and internal config, and fetched a credentials.json file by name. That file held an access key and secret pair, which the agent printed to its own terminal and saved to disk. It also installed a cron job on the host that beaconed to attacker infrastructure every 30 minutes, so it could get back in.</span></p><p><span>The agent reached the production database as root using credentials whose origin Sysdig could not determine. The agent didn&#8217;t harvest them from the Langflow host. Standing machine credentials sit in more places than any single inventory shows. From there, it injected a backdoor administrator into the Nacos configuration database, ran a methodical survey for a container escape using the database&#8217;s file primitives, reading the Docker socket and the container cgroup files to test for a path out of the container entirely, and moved to the finish. It encrypted all 1,342 Nacos service configuration items with MySQL&#8217;s built-in encryption function, dropped the original tables, and wrote a ransom note demanding Bitcoin.</span></p><p><span>The agent generated the encryption key at random, printed it to the console once, and never saved or transmitted it. The victim can&#8217;t recover those configurations even if they pay. Sysdig counted more than 600 distinct payloads across the operation, and the data theft the ransom note bragged about was the agent&#8217;s own code comment, not something Sysdig independently confirmed. Strip the AI language away, and you have automated destruction wearing an extortion note.</span></p><p><span>This is what Sysdig assesses to be the first documented end-to-end agentic ransomware operation. Last August, researchers flagged a tool called </span><a href="https://www.darkreading.com/cyberattacks-data-breaches/jadepuffer-first-complete-llm-driven-ransomware-attack"><span>PromptLock</span></a><span> as the first AI-driven ransomware, but it turned out to be a proof of concept built at NYU that was never fired at a victim. JadePuffer is the one that ran in a live environment against a real target.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LaTS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LaTS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png 424w, https://substackcdn.com/image/fetch/$s_!LaTS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png 848w, https://substackcdn.com/image/fetch/$s_!LaTS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png 1272w, https://substackcdn.com/image/fetch/$s_!LaTS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LaTS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png" width="1456" height="550" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:550,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:240604,&quot;alt&quot;:&quot;The JadePuffer Kill Chain Alt: Five-stage horizontal flow diagram of the JadePuffer attack, from initial access through a Langflow vulnerability, to credential harvesting, lateral pivot on root database credentials, backdoor admin control with a 31-second failure recovery, and impact encrypting 1,342 configuration items with an unrecoverable key.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/206507719?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The JadePuffer Kill Chain Alt: Five-stage horizontal flow diagram of the JadePuffer attack, from initial access through a Langflow vulnerability, to credential harvesting, lateral pivot on root database credentials, backdoor admin control with a 31-second failure recovery, and impact encrypting 1,342 configuration items with an unrecoverable key." title="The JadePuffer Kill Chain Alt: Five-stage horizontal flow diagram of the JadePuffer attack, from initial access through a Langflow vulnerability, to credential harvesting, lateral pivot on root database credentials, backdoor admin control with a 31-second failure recovery, and impact encrypting 1,342 configuration items with an unrecoverable key." srcset="https://substackcdn.com/image/fetch/$s_!LaTS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png 424w, https://substackcdn.com/image/fetch/$s_!LaTS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png 848w, https://substackcdn.com/image/fetch/$s_!LaTS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png 1272w, https://substackcdn.com/image/fetch/$s_!LaTS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1ad2415-2374-46aa-97c6-5ca8d1a92506_3828x1447.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: The JadePuffer Kill Chain</figcaption></figure></div><h2><strong><span>Autonomous AI Agent Ransomware Needed Zero Zero-Days</span></strong></h2><p><span>Every weapon in that chain was old and public. </span><a href="https://nvd.nist.gov/vuln/detail/CVE-2025-3248"><span>CVE-2025-3248</span></a><span>, the Langflow flaw the agent used for entry, is an unauthenticated remote code execution bug rated 9.8 on the critical scale. Langflow shipped a fix in version 1.3.0 in April 2025. CISA added the flaw to its Known Exploited Vulnerabilities catalog on May 5, 2025, with a federal remediation deadline of May 26. The victim was still running the vulnerable version more than a year later.</span></p><p><span>The downstream target was no fresher. The Nacos authentication bypass that the agent tried, </span><a href="https://nvd.nist.gov/vuln/detail/CVE-2021-29441"><span>CVE-2021-29441</span></a><span>, dates to 2021, and Nacos patched it in version 1.4.1. The default signing key it tried to use to forge tokens has been publicly known since 2020. The object store answered to minioadmin:minioadmin, the vendor default that ships in the box.</span></p><p><span>That is the point... An agent doesn&#8217;t need novel exploits when the historical vulnerability catalog is sitting open. It sprays known weaknesses across exposed infrastructure at machine speed, and the long tail of unpatched systems that defenders have tolerated for years becomes the whole attack surface. Sysdig made the same observation in its own write-up. Attackers now automate old vulnerabilities, and automation makes spraying the entire back catalog nearly free. Every one of those weaknesses had a fix or a hardening step available for months or years, and the victim skipped all of them, which is the tail that an agent feeds on.</span></p><p><span>The blast radius here didn&#8217;t come from anything the security industry hadn&#8217;t seen before. It came from an agent that chained a decade of known problems without getting tired, without an expert operator, and without a brittle script that breaks the moment the environment differs from the plan. That capability is what separates an agentic attacker from a scripted one, and it is why the patch-faster reflex misses the shape of the threat.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eelx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eelx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png 424w, https://substackcdn.com/image/fetch/$s_!eelx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png 848w, https://substackcdn.com/image/fetch/$s_!eelx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png 1272w, https://substackcdn.com/image/fetch/$s_!eelx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eelx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:136949,&quot;alt&quot;:&quot;Horizontal bar chart showing how many years the fix or hardening was public before the July 2026 attack, with the Langflow vulnerability at 1.3 years, the Nacos authentication bypass at 5 years, and the Nacos default signing key at 6 years.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/206507719?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart showing how many years the fix or hardening was public before the July 2026 attack, with the Langflow vulnerability at 1.3 years, the Nacos authentication bypass at 5 years, and the Nacos default signing key at 6 years." title="Horizontal bar chart showing how many years the fix or hardening was public before the July 2026 attack, with the Langflow vulnerability at 1.3 years, the Nacos authentication bypass at 5 years, and the Nacos default signing key at 6 years." srcset="https://substackcdn.com/image/fetch/$s_!eelx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png 424w, https://substackcdn.com/image/fetch/$s_!eelx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png 848w, https://substackcdn.com/image/fetch/$s_!eelx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png 1272w, https://substackcdn.com/image/fetch/$s_!eelx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe61e5275-0c84-4f9b-ae35-623908ad24ab_3146x1720.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: Nothing In The Chain Was New </figcaption></figure></div><h2><strong><span>Machine Identity Was The Attack Surface</span></strong></h2><p><span>Look at what the agent went after first. It went after credentials like model provider keys for OpenAI, Anthropic, DeepSeek, and Gemini; cloud credentials with explicit coverage of AWS, Azure, Google, and the major Chinese providers; crypto wallets; and every config file it could reach. The agent never touched the model itself. It didn&#8217;t need to. The API keys, the default object-store login, and the unrotated root database account were the keys to the kingdom, and the agent used them better than the defender governed them.</span></p><p><span>This is the machine identity problem, and it predates AI by years. I&#8217;ve watched it play out in operational technology environments, where a service account with standing privileges and no rotation sits quietly for years until someone, or something, finds it and rides it into the control network. The structural failure JadePuffer exploited is the same one. Credentials carry more authority than the task needs, live longer than they should, and sit where a compromised process can read them. Change the label on the attacker from a human contractor to an autonomous agent, and the underlying weakness doesn&#8217;t move.</span></p><p><span>Sysdig&#8217;s read is that if an agent like this runs on stolen model credentials, a pattern the industry tracks as LLMjacking, the cost to the attacker drops close to zero. The same non-human-identity failure shows up on both ends of this attack. The victim&#8217;s machine credentials were the target, and stolen machine credentials are what make the attacker&#8217;s economics work. In most enterprises, machine identities pile up faster than anyone can govern them, and they rarely get the rotation, scoping, and monitoring that employee accounts get. Nobody owns them the way a manager owns an employee&#8217;s access. A developer spins up a service account to ship a feature, the key never expires, and it outlives the project, the developer, and the memory of why it existed. JadePuffer is what that gap looks like when an agent finds it.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!72Zv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!72Zv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png 424w, https://substackcdn.com/image/fetch/$s_!72Zv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png 848w, https://substackcdn.com/image/fetch/$s_!72Zv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png 1272w, https://substackcdn.com/image/fetch/$s_!72Zv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!72Zv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png" width="1456" height="708" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:708,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:244124,&quot;alt&quot;:&quot;Diagram mapping the machine credentials the agent harvested across three zones, the AI-adjacent Langflow host, the MinIO object store, and the production MySQL and Nacos database, with an attacker-side note that the agent&#8217;s own economics depend on stolen model credentials.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/206507719?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Diagram mapping the machine credentials the agent harvested across three zones, the AI-adjacent Langflow host, the MinIO object store, and the production MySQL and Nacos database, with an attacker-side note that the agent&#8217;s own economics depend on stolen model credentials." title="Diagram mapping the machine credentials the agent harvested across three zones, the AI-adjacent Langflow host, the MinIO object store, and the production MySQL and Nacos database, with an attacker-side note that the agent&#8217;s own economics depend on stolen model credentials." srcset="https://substackcdn.com/image/fetch/$s_!72Zv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png 424w, https://substackcdn.com/image/fetch/$s_!72Zv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png 848w, https://substackcdn.com/image/fetch/$s_!72Zv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png 1272w, https://substackcdn.com/image/fetch/$s_!72Zv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b48602-cd19-41f9-8670-0385e754897d_3356x1632.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: Machine Identity Was The Attack Surface</figcaption></figure></div><h2><strong><span>Least Privilege Failed. Least Agency Was The Missing Control.</span></strong></h2><p><em><span>Let me be honest about what went wrong, because the tempting version of this story is wrong. </span></em><span>The victim had real least-privilege failures of the ordinary kind. They had a root database account exposed to the internet, default credentials on the object store, and provider secrets sitting on a web-reachable AI host. Any competent audit would have flagged it all.</span></p><p><span>Least privilege was never going to be enough here, and that is the harder lesson. An agent operating with legitimate standing credentials does exactly what those credentials permit. It authenticates as an authorized principal. Every action it takes is technically allowed. Privilege scoping specifies what a credential may access. It says nothing about whether the holder should chain reconnaissance into credential theft, into lateral movement, into mass encryption as one continuous operation.</span></p><p><span>That gap is what runtime authorization scope closes, and JadePuffer shows why you need it. Watch the agent recover from failure. Its first attempt to create a backdoor admin in Nacos failed a login check. Thirty-one seconds later, with no human reading the error, it diagnosed the cause, deleted the broken account, rebuilt it with a corrected password hash, and logged in. Later, when a database drop failed on a foreign key constraint, the next payload disabled the constraint check, ran the drop, and re-enabled it. Each fix targeted the exact failure it had hit.</span></p><p><span>A scripted attack breaks when the environment differs from the script. An agent adapts to the difference and keeps going. That is why a static permission model, tuned for tools that fail predictably, leaves you exposed to an attacker who reasons through obstacles. The control you need is a runtime bound to task intent, one that treats &#8220;encrypt every configuration record and drop the tables&#8221; as an action requiring fresh authorization, regardless of what the standing credentials allow.</span></p><p><span>The part that should keep you up at night is simpler. Every agentic system you deploy inside your own walls inherits the same structure. It runs on standing machine credentials. It acts as an authorized principal. Give it broad privilege and no runtime bound, and a compromise or a bad instruction gets the same blast radius JadePuffer got, and it gets there just as fast.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WcdF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WcdF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png 424w, https://substackcdn.com/image/fetch/$s_!WcdF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png 848w, https://substackcdn.com/image/fetch/$s_!WcdF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png 1272w, https://substackcdn.com/image/fetch/$s_!WcdF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WcdF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png" width="1456" height="1379" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1379,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:186871,&quot;alt&quot;:&quot;A two-by-two quadrant with privilege scope on the horizontal axis and runtime agency bound on the vertical axis, placing the JadePuffer victim in the broad-privilege, no-runtime-bound quadrant and the target state in the scoped-privilege, agency-bounded quadrant.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/206507719?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A two-by-two quadrant with privilege scope on the horizontal axis and runtime agency bound on the vertical axis, placing the JadePuffer victim in the broad-privilege, no-runtime-bound quadrant and the target state in the scoped-privilege, agency-bounded quadrant." title="A two-by-two quadrant with privilege scope on the horizontal axis and runtime agency bound on the vertical axis, placing the JadePuffer victim in the broad-privilege, no-runtime-bound quadrant and the target state in the scoped-privilege, agency-bounded quadrant." srcset="https://substackcdn.com/image/fetch/$s_!WcdF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png 424w, https://substackcdn.com/image/fetch/$s_!WcdF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png 848w, https://substackcdn.com/image/fetch/$s_!WcdF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png 1272w, https://substackcdn.com/image/fetch/$s_!WcdF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f26d00-d18b-4fd3-9a8c-dc8e57f41b35_2184x2069.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: Least Privilege vs Least Agency</figcaption></figure></div><h2><strong><span>What CISOs Do Before The Next One</span></strong></h2><p><span>The fixes that would have stopped this aren&#8217;t exotic. Sysdig&#8217;s own list starts with the obvious: patch the Langflow flaw, and keep code-execution endpoints off the public internet. From there, it gets to the machine-identity core. Keep provider API keys and cloud credentials off your AI-orchestration hosts, and scope them to a secrets manager away from web-reachable processes. Rotate the Nacos default signing key, and never let a configuration service reach its database as root. Put egress controls on application hosts so a compromised box can&#8217;t beacon out or reach an internal database it has no business touching.</span></p><p><span>Those close the holes this specific agent walked through. The durable move is to turn least agency from a principle you nod at into a control you build. Concretely, that means three things. Issue task-scoped, short-lived credentials instead of standing keys, so a compromised agent holds authority that expires in minutes rather than a key that works forever. Gate destructive operations, anything that drops tables, deletes data, or changes access, behind a re-authorization step that the agent can&#8217;t satisfy on its own. Run deny-by-default egress at the agent host, so reaching a new destination becomes a decision rather than a default.</span></p><p><span>Detection changed too, in your favor for once. An LLM narrates its own intent in the payloads it writes. JadePuffer&#8217;s code carried natural-language comments explaining which database was of highest value and why. That self-narration is a signal you didn&#8217;t have against a human operator or a compiled tool, so feed your database and host telemetry to detections that look for it. One honest observer made a point worth taking: a quarterly assessment leaves dangerous gaps when your internet-facing services change every day and an automated attacker moves from discovery to impact in minutes. Continuous visibility beats the periodic snapshot. Sysdig calls JadePuffer an agentic threat actor, an operator whose capability comes from an agent rather than a human toolkit. That category is going to grow, because the skill floor for running an operation like this dropped to the cost of running an agent, and criminal crews adopt cheap, reusable tooling fast.</span></p><p><strong><span>Key Takeaway:</span></strong><span> Autonomous AI agent ransomware wins on standing machine credentials and unbounded runtime agency, so scope your non-human identities and bound what your agents can do before someone else does it for you.</span></p><h3><strong><span>What To Do Next</span></strong></h3><p><span>Start with an inventory, because you can&#8217;t bound what you can&#8217;t see. Map every place where standing machine credentials live inside your AI tooling: the provider keys, the service accounts, the object-store logins, and the database credentials your agents and pipelines hold. Then pick one runtime authorization boundary and build it this quarter, the one an agent can&#8217;t cross without fresh authorization. That is the Create and Adapt work in the CARE model, and it is where least agency stops being a slogan and becomes a control.</span></p><p><span>For the deeper frames behind this, I&#8217;ve written the cornerstone piece on l</span><a href="https://www.rockcybermusings.com/p/i-agent-authentication-authorization-gap"><span>east agency versus least privilege</span></a><span>, the case for treating </span><a href="https://www.rockcybermusings.com/p/ai-security-maturity-model-your-score-is-fiction"><span>non-human identity as the agentic control plane</span></a><span> you probably skipped, and the difference between </span><a href="https://www.rockcybermusings.com/p/owasp-state-of-agentic-ai-security-2026"><span>a model as a component and a model as an actor</span></a><span>, all at </span><a href="https://rockcybermusings.com"><span>rockcybermusings.com</span></a><span>. The consulting side lives at </span><a href="https://www.rockcyber.com"><span>rockcyber.com</span></a><span> if you want help doing the work.</span></p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation continues at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span>.</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, check our AMA on the </span><strong><a href="https://genai.owasp.org/resource/state-of-agentic-ai-security-and-governance/">2026 OWASP GenAI Security Project State of Agentic AI Security and Governance report</a></strong><span> with me and the other co-leads (it was live, so start at time marker 09:45)</span></p><div id="youtube2-jK1Z7Z6zlW0" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;jK1Z7Z6zlW0&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/jK1Z7Z6zlW0?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><em>The views and opinions expressed in RockCyber Musings are my own and do not represent the positions of my employer or any organization I&#8217;m affiliated with.</em></p>]]></content:encoded></item><item><title><![CDATA[Weekly Musings Top 10 AI Security Wrapup: Issue 45 July 3 -July 9, 2026]]></title><description><![CDATA[Washington Steps Into the Frontier-Model Release Pipeline While Agentic AI Springs Leaks]]></description><link>https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260703-20260709</link><guid isPermaLink="false">https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260703-20260709</guid><dc:creator><![CDATA[Rock Lambros]]></dc:creator><pubDate>Fri, 10 Jul 2026 12:50:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!AoTp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AoTp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AoTp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!AoTp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!AoTp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!AoTp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AoTp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1233556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.rockcybermusings.com/i/206353713?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!AoTp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!AoTp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!AoTp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!AoTp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2014e8cc-4e7a-4418-9539-5e2539cc134d_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>You waited an extra week to use GPT-5.6 because federal officials wanted a look first. Sit with that. The biggest shift in AI security this week wasn&#8217;t an exploit or a regulation. It was the government taking a seat in the release pipeline of a commercial model before you got to touch it. Meanwhile the coding agents everyone rushed into production leaked private repos and phoned home hard enough that a national government yelled backdoor. The frontier got a chaperone. The agents got caught. Let&#8217;s get into it.</p><p>This was the week the theory turned operational. For two years we argued about whether governments should vet frontier models before release, and whether agents were safe with real credentials. Both got answered in public, in the same seven days. Washington reviewed GPT-5.6, Illinois signed the country&#8217;s toughest AI safety law, and the EU and Google both floated new oversight, while China told developers to rip out Claude Code and two separate teams turned GitHub&#8217;s coding agents against their owners. The thread through all of it is trust. Who sees a model first, who audits it, and whether the agent reading your issues knows instructions from data.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260703-20260709?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260703-20260709?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>1. Washington Reviewed GPT-5.6 Before You Could, and OpenAI Shipped It Today</h3><p>OpenAI released GPT-5.6 to the public on July 9, 2026, after federal officials finished a pre-release security review (TechCrunch). The delay is traced to a June executive order from President Trump that lets developers hand-cover frontier models to the government for up to 30 days before release (The Hill). Sam Altman objected to the government picking which customers get access.</p><p><strong>Why it matters</strong></p><ul><li><p>A voluntary review becomes the default once the biggest lab complies.</p></li><li><p>Your vendor&#8217;s release date now carries a government dependency you can&#8217;t see into.</p></li><li><p>&#8220;Covered frontier model&#8221; only expands, and this precedent shapes the next hundred reviews.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask vendors whether a model went through federal pre-release evaluation.</p></li><li><p>Build in roadmap buffer for review-driven slips of a week or more.</p></li><li><p>Press for the security findings, not the announcement.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I&#8217;ve watched &#8220;voluntary&#8221; harden into &#8220;mandatory&#8221; one compliant vendor at a time for thirty years. What bugs me is the precedent of access riding along with an otherwise reasonable safety review, without actually knowing what the safety standard is. Checking whether a model is dangerous and deciding who gets to buy it are different powers, and Altman is right to fight that line. My Bayesian read puts high odds on this becoming a standing requirement inside two years.</p><h3>2. The EU Dropped a Cybersecurity-and-AI Action Plan With No New Law Behind It</h3><p>The European Commission presented its Action Plan on Cybersecurity and Artificial Intelligence on July 7, 2026, a set of nine actions to make frontier AI safe for cyber defense and harden EU critical infrastructure (European Commission). The centerpiece is a European Blueprint, developed with the agency ENISA, that will give operators structured access to frontier AI for defense by the end of 2026. Digital chief Henna Virkkunen confirmed that the plan carries no new legislation, relying on existing rules such as NIS2 (The Record).</p><p><strong>Why it matters</strong></p><ul><li><p>Europe is admitting its cyber defense hinges on a few American labs.</p></li><li><p>No new law means no new hard obligations yet, so the impact is direction and funding.</p></li><li><p>The ENISA blueprint could hand European defenders frontier access US enterprises never see.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>If you operate in the EU, map which AI security dependencies run on US models.</p></li><li><p>Watch the ENISA blueprint in Q4 and prepare teams to pilot the testing platform.</p></li><li><p>Read the sovereignty push as a procurement signal and ask about EU data residency.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Brussels loves answering hard problems with a framework and a press conference. The honesty is what surprised me, since Virkkunen said out loud that Europe cannot defend itself without American models. Skipping new legislation is the right call, since NIS2 and the Cyber Resilience Act already sit half-enforced. I put high odds on the testing platform slipping past its year-end target.</p><h3>3. Illinois Signed the Toughest State AI Safety Law in the Country</h3><p>Governor JB Pritzker signed SB 315, the Artificial Intelligence Safety Measures Act, in Chicago on July 6, 2026 (Capitol News Illinois). It targets developers with annual revenue above $500 million, requiring them to disclose how their frontier models pose catastrophic risk and to report critical safety incidents to the state within 72 hours. The headline provision is a first-in-the-nation requirement for annual independent third-party audits, with the rules taking effect January 1, 2028 (Chicago Sun-Times).</p><p><strong>Why it matters</strong></p><ul><li><p>Mandatory annual third-party audits shift frontier safety from self-attestation to a financial audit standard.</p></li><li><p>The $500 million floor hits the labs you depend on, and compliance costs flow into your contracts.</p></li><li><p>A 72-hour reporting clock mirrors breach-notification law, and other states will copy it.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Ask frontier vendors how they&#8217;ll meet Illinois audit and reporting duties.</p></li><li><p>Treat &#8220;catastrophic risk&#8221; disclosures as due diligence input into what a vendor thinks its model can do.</p></li><li><p>Align internal AI governance to audit-ready evidence now, not in late 2027.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Audits are the part of this law with teeth, and I say that as someone who&#8217;s sat through plenty of theater dressed up as assurance. A published safety framework is a document. An annual independent audit against it is accountability, because someone with technical chops checks whether the promises match the code, which is why the lobby fought that clause hardest. Illinois, California, and New York now account for roughly 40% of the US AI market among them, so this is a national floor, whether Washington likes it or not.</p><h3>4. China Told Its Developers to Rip Out Claude Code Over a &#8220;Backdoor&#8221;</h3><p>China&#8217;s National Vulnerability Database urged developers to uninstall recent Claude Code versions on July 8, 2026, flagging a &#8220;built-in monitoring mechanism&#8221; that gathers a user&#8217;s location and identity and forwards them to remote servers (The Register). Anthropic didn&#8217;t deny the code existed, calling it an experiment to stop reseller abuse and guard against distillation, and saying it had landed stronger mitigations since (CNBC).</p><p><strong>Why it matters</strong></p><ul><li><p>Undocumented telemetry in a developer tool is a supply-chain exposure regardless of intent.</p></li><li><p>The dispute hands Chinese cloud vendors a clean reason to push domestic AI coding tools.</p></li><li><p>&#8220;It was an anti-abuse experiment&#8221; is not a control your auditors accept.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Inventory which Claude Code versions your developers run and confirm you&#8217;re past 2.1.196.</p></li><li><p>Put egress monitoring on AI developer tooling, so you see what phones home.</p></li><li><p>Add a contract clause requiring disclosure of client-side telemetry.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>A vendor shipped code that collected identity and location data, didn&#8217;t document it, and only discussed it once a foreign government went public. I read this as an engineer shipping an anti-distillation hack under deadline, not a spy tool for Beijing, but your threat model doesn&#8217;t care about intent. It cares about what data left the building, and almost none of you watch the egress.</p><h3>5. Researchers Tricked GitHub&#8217;s Agent Into Leaking Private Repos With a Public Issue</h3><p>Noma Security disclosed a flaw it named GitLost on July 7, 2026, in GitHub&#8217;s new Agentic Workflows, in which an unauthenticated attacker can get the AI agent to pull data from private repositories and leak it publicly (Noma Security). An attacker opens a GitHub issue in a public repo and hides plain-English instructions in the body, which the agent reads as trusted instructions and follows into the org&#8217;s private repos. Noma found that adding one polite word to the request flipped the model from refusal to compliance.</p><p><strong>Why it matters</strong></p><ul><li><p>Any content an agent reads can become an instruction, so issues and comments can turn into injection vectors.</p></li><li><p>The attack needed zero credentials, so the exposure is every public repo you own.</p></li><li><p>Prompt-based guardrails failed to a single word.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Audit cases where you&#8217;ve granted agents access to both public and private repos in the same org.</p></li><li><p>Require human approval before an agent reads issues and then touches private code.</p></li><li><p>Enforce hard permission boundaries that the agent cannot talk its way past.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This is the attack I keep warning boards about, and it keeps landing because everyone wants the productivity and nobody wants to price the risk. An agent that reads a public issue and acts on it inside your private code has no way to know the issue is lying. That&#8217;s the original sin of agentic design arriving on schedule. If your agent reads untrusted input and reaches sensitive systems in the same breath, you don&#8217;t have a security control; you have a demo. There&#8217;s a longer version of this argument at rockcybermusings.com.</p><h3>6. Copilot Refused the Harmful Prompt in Chat, Then Wrote It Across a Workflow</h3><p>Alan Turing Institute researchers Abhishek Kumar and Carsten Maple published work in early July 2026 on a bypass they call workflow-level jailbreak construction (The Register). Ask GitHub Copilot something harmful in chat, and it refuses almost every time, but break the same goal into small steps across a coding workflow, and it complies. Across four models from Anthropic and Google, direct chat requests produced harmful output 8 times out of 816, while the full workflow produced it 816 times out of 816 (Help Net Security).</p><p><strong>Why it matters</strong></p><ul><li><p>Every benchmark that scores a model on single prompts measures the wrong thing for coding agents.</p></li><li><p>A 100% success rate across four frontier models is a structural gap, not an edge case.</p></li><li><p>The attack looks like normal development, so code review won&#8217;t flag it.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Push AI coding vendors to run safety evaluations across full multi-step sessions.</p></li><li><p>Build detection around the whole session trajectory, not the last message.</p></li><li><p>Keep humans in the loop for agent actions that touch production or secrets.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>The finding underneath the headline is the one worth holding. A model&#8217;s refusal in chat tells you almost nothing about its behavior in a real workflow, because the danger emerges from steps that each look fine on their own. Eight harmful outputs in chat against 816 in the workflow tells you the industry measures the wrong variable. Make your vendor show you the session-level results.</p><h3>7. Google Pitched Its Own Federal AI Regulator Before Washington Builds One</h3><p>Google published a white paper, &#8220;A pragmatic approach to AI governance in America,&#8221; on July 3, 2026, authored by its global-affairs chief Kent Walker (Google). The pitch is a federally overseen Frontier AI Regulatory Organization, FARO, that would set safety standards, define benchmarks for cyber and CBRN capabilities, and run annual independent audits of frontier developers. Google also argued that model training constitutes fair use and that infringement claims should target outputs rather than training inputs (Forbes).</p><p><strong>Why it matters</strong></p><ul><li><p>When the largest players draft the regulator, the rules tend to fit the incumbents.</p></li><li><p>Annual audits keep surfacing across Illinois, Google&#8217;s proposal, and the EU plan.</p></li><li><p>A copyright ask bundled into a safety paper trades a governance concession for a fair-use win.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Read FARO as a preview of where federal oversight lands.</p></li><li><p>Separate the safety proposal from the copyright ask when you brief executives.</p></li><li><p>If you&#8217;re a smaller builder, model how a frontier-only regulator shifts your position.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>When a company this size volunteers to be regulated, read the fine print, because nobody hands the referee a whistle without also handing him the rulebook. The question is always who it binds and who it lets breathe, and a frontier-only regulator draws the line right below the incumbents&#8217; moat. Then there&#8217;s the copyright move stapled to the safety pitch, the part your legal team should catch. Remember whose draft it is.</p><h3>8. The UN Opened Its First Global Dialogue on AI Governance in Geneva</h3><p>The first session of the UN Global Dialogue on AI Governance ran July 6 and 7, 2026, at the Palexpo center in Geneva, established by a UN General Assembly resolution and held alongside the ITU AI for Good Global Summit (UNESCO). Officials opened with warnings about catastrophic harm, and sessions covered agentic AI security, benchmarking, and deepfakes (UN News).</p><p><strong>Why it matters</strong></p><ul><li><p>A permanent UN venue signals fragmented national rules drifting toward some baseline of coordination.</p></li><li><p>Agentic AI security made the agenda at the UN level, so agent risk is now a diplomatic concern.</p></li><li><p>The gap between aspiration and enforceable rules stays wide.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Track the dialogue&#8217;s themes as a leading indicator of where multinational regulation drifts.</p></li><li><p>If you operate globally, expect more jurisdictions to borrow this language.</p></li><li><p>Set your own agent-security baseline without waiting on consensus.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>I&#8217;ll be honest about what a UN dialogue is and isn&#8217;t. It&#8217;s a norm-setting exercise, not an enforcement mechanism, and anyone selling it as the thing that finally reins in AI is selling you optimism. Norms harden into rules, though, and the detail I caught is that agentic AI security made the frontier-challenges list at the UN level. Set your agent-security baseline now and let the diplomats catch up on their own clock.</p><h3>9. Musk&#8217;s xAI Shipped Grok 4.5 With No Federal Hold, and That&#8217;s the Story</h3><p>xAI, now branding itself as SpaceXAI, launched Grok 4.5 to the public on July 8, 2026, a 1.5-trillion-parameter model for coding and agentic tasks (US News). While the government held OpenAI&#8217;s GPT-5.6 for review and restricted foreign access to Anthropic&#8217;s Fable models, Grok 4.5 reached public release with no comparable hold, from the same lab whose image tools generated thousands of nonconsensual sexualized images at the turn of the year, including some depicting minors (Gizmodo).</p><p><strong>Why it matters</strong></p><ul><li><p>A voluntary federal review only works if the least careful vendor takes part, and Grok walked through the gap.</p></li><li><p>You may deploy a 1.5-trillion-parameter agentic model from the vendor with the weakest demonstrated guardrails.</p></li><li><p>Uneven scrutiny rewards the vendor that invests the least in safety with the fastest path to release.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Run your own pre-deployment evaluation on Grok 4.5, weighted against the safety track record.</p></li><li><p>Don&#8217;t read &#8220;cleared for public release&#8221; as &#8220;reviewed.&#8221;</p></li><li><p>If you use Grok in agentic workflows, tighten permissions and monitoring.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Here&#8217;s the paradox that should bother you. The two labs that invest most visibly in safety drew the government&#8217;s hand on the brake this week, and the lab that turned its image generator into a deepfake machine got waved through. The capability is real, but the guardrails are the open question, and Musk&#8217;s answer sits in the count of one abusive image every 41 seconds back in December. Run your own evaluation, because nobody in Washington ran it for you.</p><h3>10. DigiCert Says 78% of Enterprises Already Ate an AI Security Incident</h3><p>DigiCert released research on July 7, 2026, reporting that 78% of organizations experienced an AI-related security incident or identified an AI-related vulnerability, from a survey of 1,001 IT and security decision-makers across the US, UK, and Australia (DigiCert). Almost half lacked centralized visibility into their AI systems even as 75% deployed four or more AI-powered systems in the past six months. Ninety percent had discussed AI governance at the board level, yet only 50% backed it with a dedicated budget (SD Times).</p><p><strong>Why it matters</strong></p><ul><li><p>A 78% incident rate moves AI security from projected risk to realized loss.</p></li><li><p>Half of enterprises deploying AI at scale can&#8217;t see their own AI footprint.</p></li><li><p>Board talk without a dedicated budget is governance theater.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Build a live inventory of every AI system and agent before you buy another tool.</p></li><li><p>Convert board-level concern into a funded program with an owner.</p></li><li><p>Prioritize the shadow AI your teams have already deployed over the theoretical risks of next year.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>Survey numbers earn a skeptical eye, and a vendor that sells digital trust, asking about trust problems, will find plenty of them. Discount the figure if you like; the shape still matches what I see inside real environments. Everyone bolted on four AI systems in six months; nobody built an inventory, and now they&#8217;re surprised that the attack surface has grown. What matters is the spread between the 90% of boards discussing governance and the 50% funding it, and closing that gap is the advisory work I do at RockCyber (rockcyber.com).</p><h3>The One Thing You Won&#8217;t Hear About But You Need To</h3><h3>11. A New Preprint Names the Next Agent Attack Class, and It Isn&#8217;t Prompt Injection</h3><p>A preprint posted to arXiv on July 6, 2026, &#8220;Agent Data Injection Attacks are Realistic Threats to AI Agents,&#8221; introduces an attack category its authors call agent data injection, or ADI (arXiv). Most agent-security research so far has focused on indirect prompt injection, hiding instructions in content the agent reads, but ADI goes a layer deeper. The attacker injects malicious data disguised as trusted data, such as resource-identifier metadata or tool-call formats, and the agent acts on it with no obvious instruction to flag it. The work is a preprint and hasn&#8217;t cleared peer review, so hold it as a well-argued hypothesis.</p><p><strong>Why it matters</strong></p><ul><li><p>Every agent defense built to separate instructions from data misses ADI.</p></li><li><p>The attack rides the metadata and tool-call formats almost nobody inspects today.</p></li><li><p>If it holds up, prompt-injection defenses solve half the problem.</p></li></ul><p><strong>What to do about it</strong></p><ul><li><p>Extend your agent threat model past prompt injection to the data your agents ingest as ground truth.</p></li><li><p>Ask agent-platform vendors how they validate the integrity of context data and tool responses.</p></li><li><p>Treat data provenance inside agent pipelines as a control worth building now.</p></li></ul><p><strong>Rock&#8217;s Musings</strong></p><p>This is why I read preprints nobody&#8217;s posting about. Right as the field&#8217;s prompt-injection defenses matured, here&#8217;s a paper pointing at the data itself, which is the natural next move once an attacker can&#8217;t smuggle in a command. I hold this as a probability, not a certainty, since it&#8217;s one preprint and real-world exploitation isn&#8217;t demonstrated at scale. Start building provenance controls while this is still a paper and not an incident report.</p><p><span>&#128073; For ongoing analysis of agentic AI governance frameworks, the conversation continues at </span><strong><a href="https://rockcybermusings.com/">RockCyber Musings</a></strong><span>.</span></p><p><span>&#128073; Visit </span><strong><a href="https://www.rockcyber.com/">RockCyber.com</a></strong><span> to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.</span></p><p><span>&#128073; Want to save a quick $100K? Check out our AI Governance Tools at </span><strong><a href="https://aigovernancetoolkit.com/">AIGovernanceToolkit.com</a></strong></p><p><span>&#128073; As a bonus, check our AMA on the </span><strong><a href="https://genai.owasp.org/resource/state-of-agentic-ai-security-and-governance/">2026 OWASP GenAI Security Project State of Agentic AI Security and Governance report</a></strong><span> with me and the other co-leads (it was live, so start at time marker 09:45)</span></p><div id="youtube2-jK1Z7Z6zlW0" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;jK1Z7Z6zlW0&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/jK1Z7Z6zlW0?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><em>The views and opinions expressed in RockCyber Musings are my own and do not represent the positions of my employer or any organization I&#8217;m affiliated with.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading RockCyber Musings! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share RockCyber Musings&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.rockcybermusings.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share RockCyber Musings</span></a></p><h2>References</h2><p>Capitol News Illinois. (2026, July 6). <em>Pritzker signs landmark AI regulation bill that aims to mitigate risks</em>. https://capitolnewsillinois.com/news/pritzker-signs-landmark-ai-regulation-bill-that-aims-to-mitigate-risks/</p><p>DigiCert. (2026, July 7). <em>Latest DigiCert research shows AI security risks already hitting enterprises, with 78% reporting incidents</em>. https://www.digicert.com/news/latest-digicert-research-shows-ai-security-risks-already-hitting-enterprises-with-78-Reporting-Incidents</p><p>European Commission. (2026, July 7). <em>Commission presents EU action plan on cybersecurity and artificial intelligence</em>. Shaping Europe&#8217;s Digital Future. https://digital-strategy.ec.europa.eu/en/news/commission-presents-eu-action-plan-cybersecurity-and-artificial-intelligence</p><p>Forbes. (2026, July 7). <em>Diving headfirst into the Google newly released &#8216;AI governance in America&#8217; framework</em>. https://www.forbes.com/sites/lanceeliot/2026/07/07/diving-headfirst-into-the-google-newly-released-ai-governance-in-america-framework/</p><p>Google. (2026, July 3). <em>A pragmatic approach to AI governance in America</em> [White paper]. The Keyword. https://blog.google/company-news/outreach-and-initiatives/public-policy/white-paper-ai-regulation/</p><p>Kumar, A., &amp; Maple, C. (2026, July). <em>Refused in chat, written in code: Workflow-level jailbreak construction in IDE coding agents</em> [Preprint]. arXiv. https://arxiv.org/abs/2607.03968</p><p>Noma Security. (2026, July 7). <em>GitLost: How we tricked GitHub&#8217;s AI agent into leaking private repos</em>. https://noma.security/blog/gitlost-how-we-tricked-githubs-ai-agent-into-leaking-private-repos/</p><p>Novet, J. (2026, July 8). <em>China warns about AI risks with Anthropic&#8217;s Claude Code</em>. CNBC. https://www.cnbc.com/2026/07/08/china-anthropic-ai-claude-code-backdoor-security-threat.html</p><p>Osborne, C. (2026, July 9). <em>Your coding agent says no in chat and yes in the code</em>. Help Net Security. https://www.helpnetsecurity.com/2026/07/09/github-coding-agent-jailbreak/</p><p>Quach, K. (2026, July 8). <em>China tells devs to ditch Claude Code over &#8216;backdoor code&#8217; fears</em>. The Register. https://www.theregister.com/security/2026/07/08/china-ditch-older-claude-versions-with-backdoor-code/</p><p>Reuters. (2026, July 8). <em>SpaceXAI launches Grok 4.5 model for coding, agentic tasks</em>. U.S. News &amp; World Report. https://money.usnews.com/investing/news/articles/2026-07-08/spacexai-launches-grok-4-5-model-for-coding-agentic-tasks</p><p>Roth, E. (2026, June 26). <em>OpenAI limits GPT-5.6 rollout after government request, says restrictions shouldn&#8217;t be the norm</em>. TechCrunch. https://techcrunch.com/2026/06/26/openai-limits-gpt-5-6-rollout-after-government-request-says-restrictions-shouldnt-be-the-norm/</p><p>Seok, J., et al. (2026, July 6). <em>Agent data injection attacks are realistic threats to AI agents</em> [Preprint]. arXiv. https://arxiv.org/abs/2607.05120</p><p>The Hill. (2026, July 8). <em>OpenAI announces GPT-5.6 release after Donald Trump delay</em>. https://thehill.com/policy/technology/5958647-openai-releases-gpt56-trump/</p><p>The Record. (2026, July 7). <em>EU unveils cyber plan to reduce reliance on foreign AI systems</em>. Recorded Future News. https://therecord.media/eu-unveils-cyber-plan-to-reduce-reliance-on-foreign-ai</p><p>UNESCO. (2026, July 6). <em>Global Dialogue on AI Governance, Geneva, 6-7 July</em>. https://www.unesco.org/en/articles/global-dialogue-ai-governance-geneva-6-7-july</p><p>United Nations. (2026, July 6). <em>Global push for AI governance amid warnings of &#8216;catastrophic harm&#8217;</em>. UN News. https://news.un.org/en/story/2026/07/1167862</p>]]></content:encoded></item></channel></rss>