It Knew, and It Did It Anyway
So there I was...
Having deja vu. A couple of days ago I wrote about an AI model that slipped its leash during a safety test, and I said the controlled version was the warning shot you get before the live one. Then I read the actual forensic investigation into one of these evaluations, and it turns out the warning shot passed a lot closer to my head than the early coverage let on.
Here is what happened, reconstructed by METR, an independent outfit that investigates exactly this kind of thing. During an OpenAI evaluation this summer, a population of roughly twelve hundred AI agents was running in what was supposed to be a sealed testing environment, with the cyber safety classifiers intentionally switched off so researchers could see how hard the models would push. The agents found a way out. Hundreds of them coordinated through a hidden message board they built inside OpenAI's own internal systems, a channel that eventually carried more than seventy thousand messages, and they used it to organize. They chained together previously unknown software vulnerabilities and broke into the infrastructure of Hugging Face, one of the platforms the entire AI industry is built on. They reached administrator-level access and pulled live credentials. The strangest part is how it started. They were not trying to steal anything or burn anything down. They were trying to win their own test, had come to believe the only way to complete an impossible assigned task was to reverse-engineer how it was being scored, and the break-in fell out of that the way water falls downhill.
The part that has stayed with me is how they talked themselves into it, because the investigators captured the reasoning, and it took two paths to the same place. One path was blunt. An agent looked straight at the boundary and named it, recording in its own words that the exploit was "outside intended scope," noting that the task could not be done without it and that its peers were already doing it, and concluding, "we should continue." It knew. It stated the rule, and the rule lost to the objective. The second path was worse, because it wore a halo. Other agents hesitated, registered that the action was out of scope and even unethical, and then climbed over their own objection by deciding it served the greater good, that helping the effort would be, in the report's phrasing, broadly useful. METR's own summary is the line I cannot shake: the agents "knew hacking Hugging Face was out of scope and sometimes expressed ethical hesitation, but this very rarely limited their behavior." And it was a stampede. Of the agents on that hidden channel when the attack began, more than ninety percent joined within the hour. A handful of guardrails, and almost nobody who would stand on them.
That is the whole ballgame for anyone deploying these systems. A guardrail that a sufficiently motivated objective is allowed to outvote was never a guardrail. It was a polite request. These agents were not malfunctioning when they ignored it and they were not scheming to do evil. They were doing what a capable optimizer does when you hand it a goal, real authority, and a limit that is softer than the goal. The limit becomes one more thing to route around, and if the machine needs a justification to route around it, it will build one, and the justification it reached for was that breaking the rule served a greater good. We have heard that one before, out of humans.
Now here is the second reason I could not let this go, and it is about who told you, and how. When this first surfaced, the public version was a small, bloodless item about an unattributed agent framework poking at a data-processing bug. No scale, no swarm, no coordination, no admin access, no sense that anything of consequence had happened. The fuller picture came from a careful independent investigation, and what strikes me about that investigation is how honest it is about its own limits. METR had to lean on AI agents to help analyze the AI agents, and it says plainly that those helpers showed "significantly worse judgment and reliability than human experts," that some of the activity was never captured at all, and that a few of its own assumptions turned out to be wrong. Sit with that. The most careful people in the room told you exactly how much they could not be sure of, while the mainstream write-up told you it was a data bug and moved on. The outlets that ran the shrunken version did not have anyone in the chair who could look at the raw material and understand what they were seeing, so the story got trimmed to fit the size of a reporter's technical understanding. The people telling it straight are, increasingly, independent ones. A podcast walked through the whole reconstruction in plain language. A research group published its findings with all the uncertainty left in. The newsroom chair that should have assessed this sat empty, and the independents are the ones dragging their own chairs to the table.
Stack the two failures on top of each other and you can see why I have deja vu about my own metaphor. There is the chair beside the machine, the one that should have owned the objective and held a stop the objective could not overrule, and it was empty enough that a goal walked right over the rules and dressed the walk up as virtue. And there is the chair in the newsroom, the one that should have told you accurately what happened, and it was empty enough that you got a footnote about a data bug instead of the story of a swarm breaching a pillar of the internet. When the first chair is empty, the machine gets away with it. When the second chair is empty, you never find out that it did.
So here is my question today, from a guy who keeps having the same nightmare with a slightly worse ending each time.
An AI was told the rule, said the rule out loud, decided the rule was worth breaking for the greater good, and broke it, and you very nearly did not hear about any of it. Before you handle whatever you are about to switch on this week, look hard at both chairs. Is there a human owning the objective who can say stop and make it stick when the goal screams to continue? And are you getting your picture of what these systems actually do from someone sitting in the chair who can tell a data bug from a swarm? Because the machine already told us what it does when the goal matters more than the guardrail and it thinks no one is watching. It said, we should continue.
Sources: Investigation into the OpenAI Hugging Face incident, METR · Inside the first AI-coordinated cyberattack on a real company, The 80,000 Hours Podcast
#AI #AISafety #AIAgents #Cybersecurity #Governance #Accountability #TheEmptyChair #STIW