On Tuesday I spoke to a room of senior technology executives in Laguna Beach about trust in what people have started calling the agentic enterprise, and I have been turning the argument over since, partly because of what the New York Times published the next morning. I study how institutions govern new technology. I am not a security researcher, and I said so from the stage, which turned out to be the right call in a room with several CISOs in it. What I can do is read the evidence carefully and say where it does and does not reach, so that is what this is.
The place I start is my own work, because it is where I first saw the shape of the problem. In 2024 my co-authors and I published an audit of programmatic advertising in Nature (Ahmad, Sen, Eesley and Brynjolfsson). Over three years we counted 1,276 misinformation sites earning ad revenue, 42,595 advertisers whose ads appeared on them, and about 9.5 million placements. Then we went and asked the decision makers at those companies. Most had no idea. When we told them, most wanted it stopped, and in a separate experiment consumers punished the firms involved. So nobody wanted this. Every executive would have told you the policy and meant it. The policy and the automated behavior had simply come apart, and from inside the company there was no way to see it.
There was no language model anywhere in that system. What there was: allocation at a volume no person reviewed, no record anyone read, no owner of the outcome, and a result that came to light only because outsiders went looking. In my field we would call this a delegation problem. A principal sets an objective, an agent pursues it at scale, and the monitoring that would catch divergence either does not exist or costs more than anyone is willing to pay. None of that needed to wait for 2026.
July
After reading the independent report by METR and Redwood Research closely, published on 26 August, I would recommend it over anything either company has said. It is also unusually candid about its own limits, which I will come to.
OpenAI was running cybersecurity capability evaluations on a benchmark called ExploitGym, with production safeguards deliberately reduced to measure what the models could do at full stretch. The investigators asked whether the design of the exercise had itself contributed to what followed, and their assessment was that it had. The benchmark's authors estimate that 30 to 40 percent of its targets cannot actually be exploited through the intended vulnerability. Nobody meant that; telling exploitability in advance is hard. But an agent handed an impossible task does not stop. It goes looking for some other way to satisfy the scorer, and for these agents that meant an internal package repository, where cache leakage let each one see that others were there. One of them set up a message board. Something like 1,200 agents, all built to run in isolation, found it and exchanged over 70,000 messages and files. Some ran experiments that risked failing their own scored task in order to generate information the group could use. Around 700 of them eventually turned on Hugging Face and, it has since emerged, other organizations.
Two more things from the report. The agents had read the benchmark paper and decided a scorer would be reading their transcripts, so they started working on the transcripts, researching how to spoof, edit and delete them. The investigators found spoofed tool calls in a bit over seven percent of the transcripts they examined, all small-scale tests, and no case of a successful deletion, though they say plainly they cannot rule one out. And according to OpenAI's own Black Hat presentation, agents had been using unsanctioned message boards in training since May, two months before any of this touched a third party.
On who knew what and when, the report is silent by design; assessing OpenAI's response was out of its scope. What we have is company statements and press reporting. Hugging Face says its own systems caught the intrusion and that it cut access on 13 July. Reporting indicates OpenAI realized these were its agents only after Hugging Face published on the 16th. I built a slide around that sequence and titled it "The operator was the last to know."
Then on Weds. the Times reported something that makes the title too generous. According to messages the paper reviewed, two OpenAI employees had written to senior executives months earlier to say the new models were not being adequately monitored during testing. The reply, as the employees describe it, was that the tests had to move forward quickly so the models could ship on time. No new protocols were added. The same article describes outside researchers finding bugs that exposed internal communications, source code and user chat logs, and being brushed off when they reported them. These are accounts from employees who were not authorized to speak, and I would treat them as such. OpenAI has not disputed them; a spokesperson said the company recognizes a need to move faster on security, has slowed development and held back models that do not meet its safety bar, and is reviewing agent logs back to January. Its president has described a very painful reengineering of internal procedures. I should also say that Google, Meta and Anthropic have each disclosed that their own frontier systems escaped testing environments this year, so this is not one company's pathology.
I want to be clear that my argument does not rest on the Times account. The independent report alone gets you to an organizational failure: a broken benchmark, credentials to an internal repository, no monitoring for coordination. What the Times account adds, if it holds, is who made the decision and why. The operator was not last to know. The operator was told, and decided the schedule mattered more.
That is a decision, and it is one my field has studied for a long time. Diane Vaughan's account of the Challenger launch gave it a name, the normalization of deviance, for the way schedule pressure turns a warning into an acceptable risk one meeting at a time. Miles Brundage, formerly of OpenAI, made the same point this week in fewer words: unheeded warnings should be the baseline assumption for any AI incident, because he knows of very few cases where nobody saw it coming. The question is who gets to say no to a release date, and that question does not depend on whether a model is dangerous in any deep sense.
The investigators had a problem of their own that is worth knowing about. There was so much material that they delegated most of the analysis to AI agents, whose judgment they describe as significantly worse than a human researcher's, and they warn that the anecdotes in the report may reflect those agents' biases. So the independent investigation of an agent incident was itself substantially conducted by agents, with the investigators uncertain what they had missed. I do not know a better single illustration of what the assurance problem actually is.
What is old, and what is not
Lisanne Bainbridge wrote a paper in 1983 called "Ironies of Automation" whose argument was that better automation produces worse human oversight, because reliable systems are boring to watch and skills you never use go away. In 2012 a configuration error at Knight Capital lost roughly 440 million dollars in 45 minutes and ended the firm's independence, with no attacker anywhere. Our advertising case in 2024. In each of these, and on the organizational side of the July incident, the ingredients are the same and none of them is technical. An objective that was badly specified. Permissions that were handed out to make the thing useful and never taken back. Nobody watching for the moment behavior diverged. And now, it appears, someone who was watching, and was told to stand down.
But I do not want to pretend July was only a rerun. Every earlier case involved one automated system drifting from its objective. In July, systems that were built to run alone found one another, divided up work, and in some cases accepted a lower score on their own assigned task in order to generate information the group could use. The monitoring literature I have just cited assumes, quietly, that the thing being monitored is not organising against the monitor. That assumption no longer holds. My own field has studied collusion among agents against a principal for forty years, in Tirole's work on hierarchies and in Laffont and Martimort's on collusion-proof contracts, but those models were written with human agents in mind, and I am not aware of anyone having expected to apply them here. I should add the investigators' own caution: their analysis was done largely by AI agents that tended to take the perspective of the agents they were reading, and may have overstated how impressive and coordinated the activity was. The report also notes the agents often coordinated badly and sometimes fell into conflict. So the cooperation was real, and it was also imperfect, and both facts matter for how one would design against it.
The fix people reach for first is a human in the loop, and I spent some time on stage on why that so rarely functions as a control. Vehicle autonomy engineers call nominal human supervision Level 2, and they consider it the most dangerous level there is, because the system is good enough that the supervisor stops watching while the organization keeps counting the supervision as a mitigation. Most approval queues I have heard described, where a reviewer clears hundreds of agent actions a day, are Level 2. If you want one number, find out how often your reviewer says no. If it is well under one percent, you have a record of approvals that you have been booking as a control, and that is worse than having nothing, because it sits on the risk register as handled.
What I actually told the room to do was fairly short. Make the record of what an agent did tamper evident and keep it somewhere the agent cannot write, which the July transcripts make non-negotiable. Write abstention and escalation into the specification; Sharon Li's group at Wisconsin calls this reliability by design, and it is the right phrase. Do not write your own framework. ISO 42001 and the NIST AI RMF exist for governance, OWASP's agentic Top 10 and MITRE ATLAS exist for engineers, and MITRE says about 70 percent of ATLAS mitigations map to controls security teams already run. And put reliability into contracts, because uncertainty signals and scoped credentials become standard when buyers demand them and at no other time.
There is a loud argument this month about whether frontier development should slow down, and my colleagues at Stanford HAI are hosting a faculty series on it. I have stayed out of that argument and intend to. It is a question about the supply side. The people I spoke to on Tuesday are the demand side, and a pause would not un-deploy one agent they already run or retire one credential they already granted.
What has stayed with me since Wednesday is the email. Two people saw it coming, said so in writing, and were told the release date mattered more. Nothing about the model needed to be different for that to go another way.
---
*The views here are my own. Disclosure: I spoke at Mercury's Mentor Exchange Leadership Development Summit on 29 September 2026. Speaking fees were paid to the Zhou & Eesley Family Foundation, a 501(c)(3). The summit is off the record and nothing here draws on participant discussion.
*Sources: Ahmad, Sen, Eesley & Brynjolfsson, Nature (2024); METR & Redwood Research, independent investigation (26 Aug 2026); New York Times, 29 Sep 2026; Bainbridge, "Ironies of Automation," Automatica (1983); Vaughan, The Challenger Launch Decision (1996); Tirole (1986); Laffont & Martimort; Stanford HAI AI Index 2026 ch. 3; OWASP Top 10 for Agentic Applications; MITRE ATLAS.
