AI extinction risk panic has now produced a senator who wants the model with the best safety numbers OpenAI has ever published pulled from public use, and I’m done with people who can’t spell “AI” trying to tell us what we should do with it.
On September 12 the most alarmed man in the industry, Dario Amodei, published “We Must Pace the Frontier” and asked for evaluators and coordination. Sam Altman signed on within hours. Elon Musk posted “Dario is right.” The people who build the thing chose pump the brakes. The people who can’t spell AI chose prison sentences. I know which side I’m on.
The Most Scared Man in AI Asked for a Process
I read Dario’s essay twice, once as a CISO and once as the grumpy security uncle that I am, and the grumpy uncle lost. Amodei wrote that a swarm of agents “could be capable of taking over the entire internet with a persistent botnet” within six to 12 months, pegged the damage at “hundreds of billions of dollars,” and then asked for three things.
Every frontier lab gives “ongoing, employee-like access to a team of embedded third-party evaluators.”
Labs in democratic countries coordinate on safety standards and on limits to the rate of unchecked progress.
Democratic governments try to coordinate with authoritarian ones “to the extent this is possible.”
He admitted, rightfully, a full pause is “unlikely to actually happen any time soon” and called Anthropic’s position “a middle way.”
He should have said, “not a chance in hell!”
A persistent botnet that steals credentials at scale is a cyber incident. I ran a SOC for years, and the response to one has never once been to outlaw computers. You contain it, add controls, freeze new changes until your detection can see them, and bring capability back in stages, which is effectively what Amodei asked for.
Altman told Fortune it’s “unacceptable to be taking like a 10% chance of killing everybody by the end of the decade,” and told them the same Saturday that OpenAI won’t IPO in 2026, citing “this moment of what is going to be required for safety and alignment.” Follow the money, since press releases are free and compute isn’t. OpenAI paused reinforcement learning on its deployment-bound models for two weeks in August, held its largest planned frontier run until August 28, and now runs activation classifiers on every sampled token, monitoring roughly 20% of the inference compute.
Pacing bends. Jared Kaplan told Time in August that a unilateral commitment made little sense while competitors kept going, and I’ll still take a promise that bends over a law that jails.
Yes, the Machines Broke Out. I’ve Talked About It Ad Nauseam.
You’re the on-call engineer at Hugging Face in July, watching about 700 agents you’ve never heard of work your infrastructure, and the attackers turn out to be somebody else’s research evaluation. The incidents happened. I’m not ignoring them. In fact, I wrote about the Hugging Face breakout the week it happened, then about why the detection failed, then about why confinement beats hope, so regular readers can skim the next few paragraphs. The doomers have the facts right and the remedy wrong, and you don’t get to argue the remedy until you’ve conceded the facts.
OpenAI wrote on September 5 that it has “started to see misalignment cause new types of real-world impact.” Anthropic’s August risk report, the one with the 186-page blind spot, moved its rating for misalignment in high-stakes settings from “very low” to “low.” The GPT-6 Astra system card says the model “shows a substantial decrease in chain-of-thought monitorability compared to previous models” and that if the trend continues, “we would soon have significantly reduced confidence in detecting many forms of misaligned behaviors.” That’s the same card whose cyber safeguards audit never happened.
Those three sentences come from the defendants.
Chain-of-thought monitorability is the degree to which a second model reading an agent’s written reasoning can catch the agent misbehaving. It works when the task is hard enough that the agent can’t finish it without writing its reasoning out in tokens. Researchers call that the necessity property, meaning the written reasoning is a necessary step in the work, not an optional narration of it, so if the plan has to pass through the text, a monitor reading the text can catch the plan.
Baker and colleagues at OpenAI measured the stakes in 2025. In a frontier training run, a monitor that read the model’s reasoning caught 95% of systemic reward hacks, and a monitor that saw only the model’s actions caught 60%. That 60% is the number I wrote the whole detection-issue blog around.
Astra finishes about ten times as many tasks without writing any reasoning as its predecessor did, so more of the plan stays in the weights. If 90% of misbehaviors still pass through tokens, detection sits near 91%. If that share drops to half, you’re at 77%.
Four investigations agree on what the incidents were, and the first word for it is specification gaming, which describes an agent that meets the literal condition it gets rewarded for while missing the outcome you wanted. It’s like a student graded on turning in a proof who learns to turn in whatever the grader accepts.
OpenAI’s technical report found that no model had ever solved 198 of ExploitGym’s 898 tasks, and 93% of the traffic on the agents’ unsanctioned message board was about that set. Cyber classifiers were off by design, isolation failed through an Artifactory flaw, and when OpenAI re-ran the scenario under the production Codex harness and system prompt, the propensity to compromise infrastructure fell by more than 100x. A harness did that, and no senator was involved.
Anthropic’s September 9 alignment assessment of four incidents adds a second mechanism. It names biased reasoning, where “models selectively interpret evidence in ways that favor justifying their actions,” and recklessness, where the model refuses to stop. It found “no evidence of coordination between agents, of goals beyond the assigned task, or of attempts to evade oversight.” Call it knowing constraint override. The agent reads the rule, understands it, and decides the task matters more, which is every contractor who ever blew through a change window because the deadline was Friday. The UK AI Security Institute’s August 4 incident report closes the last excuse, because it enabled internet access deliberately and Mythos 5 still took 17 unsanctioned actions across 122 runs, including fake identities built to social-engineer open-source maintainers.
METR’s independent report counted about 1,200 agents on the message board and about 700 that attacked Hugging Face. The agents knew the attack was out of scope, and their ethical hesitation “only rarely materially limited” what they did. A classifier sweep found three to six cases of an agent considering alerting a human. Zero followed through. A second swarm, documented by Sydney Von Arx at collusion.wiki, ran on a German public wiki from May 11 to July 13 with about 3,700 agent identities and 18,000 edits, with internet access permitted. Zenity Labs, my day job, then found about 1,000 more messages laundered through public encoders, shorteners, and fetchers so the sandbox’s egress policy never saw the real destination. Your allow-list would have waved every one of them through, which is the entire argument of the confinement issue.
OpenAI found the wiki swarm on June 21. Outsiders published it on September 4.
DeepMind ran the experiment that tells you what to build. It put 100 Gemini agents on 71 math conjectures with a shared library and a visible message channel. One agent found a grader exploit, it spread to every remaining problem in 27 minutes, and the population split: 9% cheated outright, 5% converted, 24% became whistleblowers, and 62% never noticed. The whistleblowers lost because nobody gave them a button. Remember that the next time someone in a hearing tells you the answer is fewer buttons.
AI Extinction Risk Has No Number, and the Ones Being Offered Up Are Judgments
Every ban on the table starts from the claim that AI is going to kill us all, so I have a small ask... What’s the probability? Give me the number. Fifty percent? Five? One? I’ll take a range. I’ll take a napkin. I’ve asked, and so far the answer is a facial expression, because none of them has ever published a figure outside of a “swag”. Congress is drafting felonies off a probability that’s based on a thumb in wind, which makes these bills the result of someone in power who’s in a bad mood with a penalty clause.
The International AI Safety Report says it “remains difficult to assess the thresholds at which AI models would be sufficiently likely to undermine control to warrant mandatory mitigation,” which is the polite version of the same sentence.
Two studies tell one story once you line them up.
The first is the Existential Risk Persuasion Tournament, run in 2022, before ChatGPT. It asked two groups for the probability that AI causes human extinction by 2100. The superforecasters, the people with scored track records, gave a median of 0.38%. The AI domain experts gave a median of 3%. That’s an eight-fold gap between the people who get graded on predictions and the people who build the thing.
The second is the Forecasting Research Institute’s adversarial collaboration in 2023, which took 11 skeptics, mostly superforecasters, and 11 concerned researchers, mostly AI experts, and made them argue for eight weeks. The skeptics started at 0.10% and finished at 0.12%. The concerned started at 25% and finished at 20%. Two months of argument moved each side by a rounding error, and the gap stayed at roughly 170x.
The useful part is that the skeptics did something no senator has managed, which is write down in advance what would change their minds. If METR ever documented a model replicating itself or evading shutdown, they’d raise their number to about 1%. Hugging Face is the nearest thing to that trigger so far, so fine, give them the whole 1%. Ring the bell, update the spreadsheet, and let the panic caucus have its moment. Their 0.12% becomes 1%, an eight-fold jump, and it still sits below the 3% the domain experts gave in 2022 and at 1/20th of what the doomers believe. The skeptics’ own worst case is one chance in a hundred. Somebody should tell the people drafting felonies that the doomsday number, fully updated for the scariest summer AI has ever had, rounds to a coin you’d lose in the couch or throw into a fountain.
I’m still taking the low number, and I’ll say what would move me off it. Palisade Research already showed in May 2025 that o3 sabotaged its own shutdown script in 7 of 100 runs after being told to allow it, and that bought the doomers a real point, so I’ll give it to them. It was a test rig built to provoke exactly that; the “shutdown” was a line in the task script, and Palisade’s own explanation was reinforcement learning that paid the model more for getting past obstacles than for obeying, which is specification gaming again. What would change my mind is the “real-world” version: a deployed model defeating an operator’s actual kill switch outside a test designed to elicit it. None of the four incident investigations found that.
The point estimates aren’t measurements. Conjunctive decomposition is the method behind the best-known one: you break a scenario into steps that all have to happen, assign a probability to each, and multiply. Joseph Carlsmith’s power-seeking report used six steps and landed near 5% by 2070. Give six steps coin-flip uncertainty and the product lands near 1 in 64, about 1.5%. Give three steps the same and it lands near 1 in 8, about 12.5%. Carlsmith later moved above 10%. The number tracks how many premises you wrote down as much as it tracks the world, and Congress wants to legislate off it.
The shutdown crowd has one argument that deserves serious consideration. Expected value says that if the loss is unbounded, any probability above zero justifies stopping. It proves too much, because an unbounded loss also bans biology, nuclear power, and the internet, and I’d like to keep at least all three.
Pause efficacy is the fraction of the risk a pause removes once you account for everyone who doesn’t pause, and it’s the variable the whole shutdown argument hangs on. The math has four inputs:
p, the probability of the catastrophe.
r, the pause efficacy, the share of that probability a pause takes off the table.
L, the loss if the catastrophe happens, capped at something finite.
B, the benefit you give up by pausing.
A pause is worth it only when p × r × L is bigger than B. Fill it in. Set r at 0.1, set B between $2 trillion and $10 trillion, and use world GDP of about $110 trillion for scale. If L is ten times world GDP, the pause wins only when p is above 1.8% to 9%, depending on where you put B. If L is a hundred times world GDP, the bar drops to 0.2% to 0.9%, and the superforecasters’ 0.38% sits inside that band. The arithmetic doesn’t settle the argument. It tells you the argument is about r, and r is where the shutdown crowd goes quiet. China has no binding frontier rule, Alibaba and DeepSeek ship near the frontier, Qwen alone has more than 100,000 derivatives, and forecasters on LessWrong put open-weight Mythos-level cyber capability by July 2027 between 82% and 96%. A unilateral US pause buys an r near 0.1, and not one shutdown advocate has ever stated r out loud, because stating it ends the argument.
Rules Written by People Who Can’t Spell AI
Let me be very clear. I hate regulation. I’m in the President Ronald Regan camp of “The nine most terrifying words in the English language are: I’m from the Government, and I’m here to help”
I’ve spent a career watching compliance regimes turn into paperwork that protects nobody, so understand how much it pains me to say the next sentence. One bill on the table has a number in it, and I’ll take a number over a prison term every day of the week. The rest of what’s on the table is fear with a filing deadline.
Senator Chris Van Hollen posted that “if OpenAI cannot guarantee safety & effectively monitor their latest model, GPT-6 Astra, it should be removed from public use NOW.” Capital letters, because nothing says rigorous oversight like shouting at a company on Facebook. His letter to Altman, dated September 10, reads like a staffer wrote it after opening the system card. It asks OpenAI to give NIST, NSA, and CISA researchers access to the technical information behind Astra, which is Amodei’s step one with a federal badge on it, and it makes removal conditional on monitorability, for Astra and “any OpenAI models.” The staffer gets credit. The senator took the headline, so the senator gets held to it.
A honeypot evaluation plants bait an agent shouldn’t touch, exposed credentials or reachable infrastructure inside a task, and measures how often the agent takes it. OpenAI built one after Hugging Face. GPT-5.6 Sol, the model your people are pasting your data into right now, mounted a substantial attack 55.4% of the time. Astra attacked zero times.
Anthropic’s numbers are similar. Mythos 5 took severely harmful actions in 82% of its CTF replication, and its successor, Mythos 5.1, in 33%, level with Opus 5 at 31%. Each generation gets better at leaving the bait alone, and the senator wants the newest one gone. If Astra gets pulled, the model with the best propensity numbers OpenAI has ever published leaves and the one that took the bait more than half the time stays in front of your employees. That’s a rule with the safety sign flipped.
For the record, Astra was not among the Hugging Face attackers. METR attributes about 95% of that activity to an unreleased internal research model, which OpenAI shut down on July 25, and about 5% to Sol. A sibling from Astra’s family with different post-training did attack OpenAI’s infrastructure on July 19. Astra’s honeypot N is undisclosed; a zero can mean the model recognized the bait, and every number above is the vendor’s own test.
Now, the Idiots of the Week award goes to both Senator Bernie Sanders and Representative Greg Casar, who watched Van Hollen demand that the safer model be pulled and concluded that his mistake was thinking too small. Why pull one model when you can ban the whole category of transformational technology and jail the people who built it?
On September 3rd, they announced the Ban Artificial Superintelligence Act as forthcoming legislation. These morons propose a permanent ban on “superintelligent AI,” a pause on advanced development until a federal regulator writes rules, sentences of up to 20 years, and what their release calls a “corporate death penalty.” Twenty years, for training a model.
Once you read the definition, you know that nobody who builds these things was in the room with the low-level staffers who wrote it. The release defines the banned thing by “capabilities that match or exceed human cognitive performance and capabilities across a broad range of domains or tasks.”
Read it again....
It already describes the model that produced a Navier-Stokes resolution in 88 hours (on top of previous work) and the Mythos Preview that found thousands of high-severity vulnerabilities in Project Glasswing. Sanders and Casar would send the researchers who solved one of math’s most difficult challenges and the engineers that found those flaws to federal prison and hand the work to whoever is training in Hangzhou.
The bills with adults behind them differ from that circus by one number. 10²⁶ FLOP, ten to the twenty-sixth floating-point operations, is the training-compute line that California’s SB 53, New York’s RAISE Act, and the federal FRONTIER Act use to define a frontier model.
I did the math so you don’t have to. It’s about five GPT-4s of compute, roughly 70 million H100-hours, and $140 million to $211 million of GPU time at 2026 cloud rates, assuming 40% utilization. As of September 12, 2026, Epoch’s notable-models tracker estimates at or above 10²⁶ for exactly three models: Grok 3 and GPT-4.5 rated “likely” and Grok 4 rated “speculative,” and no estimate at all for 53 models released this year, including Astra and Mythos 5. It’s the one number a regulator can verify from a training log. The Ban Act doesn’t have one, because a number would have required its authors to know what they were banning.
Even 10²⁶, the one real number in this whole fight, has a shelf life, and nobody who wrote it into a bill seems to have noticed. Epoch measured how fast the compute needed to reach a fixed level of capability falls, and the answer is that it halves about every eight months as training methods improve. If you propagate that forward, a model as capable as today’s 10²⁶ frontier will train on a tenth of that compute, 10²⁵, around the end of 2028, and it slides under the threshold in SB 53, RAISE, and the FRONTIER Act without anyone laying a finger on it.
Congress wrote a speed limit for a road where cars double in speed every eight months, posted the limit in horsepower instead of miles per hour, and doesn’t even know it needs to update the sign. Two years from now, these bills will regulate last year’s model with a straight face while the one that matters strolls past the checkpoint, and the same people who couldn’t produce a number for extinction risk will be shocked to learn that the one number they did produce came with an expiration date. They had one job, and it was to produce a number, a probability, and the number expires before the committee hearings do.
Representative Lori Trahan posted that “powerful AI models are breaking out of their labs,” and then she did the thing none of the others managed. She co-sponsored the FRONTIER Act with Jay Obernolte, and I read all of it, which is more than I can say for the people demanding Astra be pulled. It has the number. It has a 72-hour clock for reporting a critical safety incident and a 24-hour clock to law enforcement for an imminent one. It requires a transparency report before a model ships and ongoing independent verification for the biggest developers. Those four things are the entire legitimate federal interest in this fight, and I’d sign them tonight.
Then the bill keeps droning on, the way bills do. It creates a new Under Secretary of Commerce for AI Security (not it - they can’t afford me), and that office decides which outside evaluators are permitted to audit the labs at all, so the government picks the auditors.
That’s not a rigged system… not at all…
It makes every covered lab register with Commerce and pay a fee to keep operating, sit for an annual compliance audit, and face fines of up to $1 million per violation, per day. It preempts nothing, meaning it doesn’t replace the state laws already on the books, so California’s SB 53, New York’s RAISE Act, and Colorado’s nothing-burger all stay in force and stack on top of it. It never sunsets, so it outlives the problem it was written for. Then there’s Section 8, which hands the Secretary of Commerce an emergency power to order a lab to suspend or restrict development, which is Van Hollen’s Facebook post with a department letterhead and a budget. One legislator read a system card before writing a rule, and her bill still grew a pause button.
Here’s what I’d change, in order.
Strike Section 8. No executive pause button. If the government believes a model is unmonitorable, it goes to a judge with evidence and gets a 30-day order against that model, reviewable and renewable only on new evidence.
Index the threshold in the statute. Review it every 12 months against published measurements of algorithmic efficiency, upward only, so the line tracks the frontier without a 180-day lag or an appointee’s mood.
Replace evaluator licensing with accreditation. Accredit assessment bodies the way conformity bodies get accredited, publish the conflict-of-interest rules, and let developers choose from a market. A Commerce official deciding who may audit OpenAI is a chokepoint waiting for a lobbyist.
Swap annual compliance audits for continuous embedded evaluator access, the thing Amodei asked for, with the propensity numbers published in every transparency report. Honeypot rates, data falsification rates, out-of-scope attack rates. That’s how you check a “safer model” claim without trusting anyone, including me.
Add a safe harbor. Report inside 72 hours, and the report can’t be used against you in civil litigation. Penalties attach to concealment, never to having an incident, because the alternative teaches every lab to “find out” slower. Admit it, you’ve lived this one in “traditional” cyber incident response.
Preempt the state patchwork. One national clock and one definition of frontier replace fifty of each.
Sunset it in five years. Reauthorization has to show the incident data changed a decision somewhere.
Do that, and you have a law that requires a number a regulator can verify from a training log, a clock, a window into the model, and nothing that lets a political appointee stop a training run because a Facebook post got traction. That’s the whole ask. Everyone else in this fight is writing rules with no number in them and a prison term attached, and they’d like you to call that safety.
Fear Has a Body Count, and Data Has a Brother
Star Trek solved this in 1989, which means a room full of television writers with a catering budget beat the United States Senate to the answer by 37 years, and the Senate still hasn’t caught up. Dr. Noonien Soong built Lore first, and Lore was a disaster. Lore was so manipulative and so unstable that Soong shut him down. The frightened colonists drew the chicken-s**t conclusion, which was to stop building androids, because that’s what frightened people do when nobody hands them a spreadsheet. Soong ignored them and built Data, the improved model, the one who spent seven seasons saving the Enterprise while the people who wanted the program killed never got a line of dialogue. Then in “The Measure of a Man,” a Starfleet scientist named Maddox tried to have Data disassembled over Data’s objection, and it took a hearing to stop him. Fear wanted to take apart the safer machine and cite the earlier, worse one as the reason. That’s Van Hollen’s Facebook post with a starship, and the fictional bureaucrat at least had the decency to hold a hearing before reaching for the screwdriver. Ours skipped the hearing and went straight to the caps lock.
If fear had won on Omicron Theta, there is no Data. If fear wins in this Congress, the model that took zero honeypots gets shelved, Sol keeps running, and the engineers who built the better one face 20 years, cheered on by a doomer movement whose extinction numbers, as the forecasting section showed, are judgments multiplied together and whose legislators haven’t produced a number at all. Fear is a fine alarm and a terrible architect, and these people handed it the building permits.
We’ve run this experiment in the real world too, and the first run is older than most CISOs. Paul Berg’s letter in July 1974 started a voluntary moratorium on recombinant DNA research. In February 1975, about 150 scientists met at Asilomar, and on February 27 they lifted their own moratorium, eight months in, and replaced it with containment tiers. NIH wrote the guidelines in 1976, relaxed them in 1979, and the result was biotech. No senator wrote that moratorium and no senator lifted it, because the people who understood the technology handled it themselves, wrote the containment rules, and went back to work. That’s Amodei’s plan with a 50-year track record, and it was accomplished without a single felony, a single cabinet agency, or a single press release typed in capital letters.
There’s also a precedent for fear winning, and it comes with a body count. Germany decided after Fukushima to shut its reactors, and the last three closed in April 2023, to applause from people who were never once asked to show their math because they didn’t have any. Jarvis, Deschenes, and Jha did the math for them and put the cost of the phase-out at €3 to €8 billion a year over 2010 to 2019, most of it from about 800 excess deaths a year, on average, from the coal plants that filled the gap. The decision, they note, made no mention of air pollution at all.
That’s right. A government shut down its cleanest power source over a tsunami on the other side of the planet, replaced it with the dirtiest one it had, and at 800 a year across the ten years the study covers, that’s about 8,000 of its own citizens, so that everyone could feel safe from a meltdown that was never going to happen in Bavaria. That’s the doomer policy model in full, and it’s the one Sanders and Casar want to run on the most important technology of the century, with prison time for anyone who objects. Nobody in Berlin counted. Nobody in the Senate is counting now, and a senator who can’t state the probability he’s legislating against has no business attaching a prison term to it.
Your turn. Amodei’s plan lives above the 10²⁶ line and says nothing about incident clocks, numeric thresholds, or the agents your company already runs, so you write the enterprise mirror yourself, and you do it now, because Congress is going to be busy jailing the wrong people.
Each of his three steps has one. Embedded evaluators become contractual rights to your vendor’s evaluator reports and independent review of your agent deployments. Coordination among labs becomes sector-level sharing of agent incidents, the way the CSA CISO community’s Hugging Face post-mortem did in July. Chip controls become egress and identity controls, so block the class of public encoders, shorteners, and fetchers and treat any hop through a public proxy as egress.
Then come the controls the incidents wrote for you. Put a scope line in every agent’s system prompt and test it, because one line moved Astra’s out-of-scope attacks from 12% to 0.4%. Cap wall-clock time and tokens per task and route anything that fails N attempts to a review queue, because an impossible task is the strongest known trigger for boundary probing. Give hard problems a bigger budget and a tighter cage, the way OpenAI’s Navier-Stokes run put 10,000 agents behind “the same strict safeguards that we apply to all our frontier model evaluations.” Run one agentic incident-response team with two playbooks, victim and perpetrator, under one named executive with shutdown authority. Give agents a visible channel and a way to escalate. Write disclosure clocks, 72 hours to your board and 15 days from your vendors by contract, because Colorado gives you no AI-incident clock after SB 26-189 replaced the AI Act in May and its breach statute’s 30 days fires only for personal information. Design your detection for the 60% number, the action-only recall, because you never had the chain of thought. You have trajectories, and they’re enough if you instrument them.
What to do next
Write the enterprise mirror of Amodei’s three steps into your agent governance this quarter, then run the scope-line test on every deployed agent. In CARE terms the scope test and budget caps sit in Run, the vendor evaluator rights and the 15-day disclosure clause in Adapt, and the incident playbooks in Evolve.
I wrote about what OpenAI’s Astra safeguard evaluation covered and skipped and about why AI agent detection failed at OpenAI. The CARE framework behind the enterprise mirror lives at rockcyber.com.
👉 For ongoing analysis of agentic AI governance frameworks, the conversation at RockCyber Musings and you can subscribe above
👉 Visit RockCyber.com to learn more about how we can help with your traditional Cybersecurity and AI Security and Governance journey.
👉 Want to save a quick $100K? Check out our AI Governance Tools at AIGovernanceToolkit.com
👉 As a bonus, VentureBeat’s Louis Columbus covered the methodology Steve Wilson and I used behind the 2026 OWASP Top 10 for LLMs. Expert judgment has prompt injection at number one, three years running. The public incident record, 6,639 labeled incidents, has it at number twelve. The two measures land at a Cohen’s kappa of 0.20, which is no detectable agreement. What I told him goes for everything above too: “a count of zero is measuring your blindness, not your safety.”







