The Teacher's Book
So there I was...
In 4th grade. It was the bicentennial, and somewhere between Schoolhouse Rock and Pet Rocks, I needed to learn my multiplication tables. Mr. H, my teacher, looked like an extra from Boogie Nights. Thick moustache, silk shirt, prolific chest hair. Gross. And he fundamentally misunderstood 4th graders. We could not be trusted. Silly man.
So, when he announced we could use the "teacher's book" to check our math homework answers and self-grade and self-correct... well, I became a math genius. My little 4th grade brain did not perceive consequences. Like punching the school bully in the face would land me in the principal's office and not carried aloft by my classmates like a returning war hero. Or the consequences of a year of not learning any math, and the humiliation of a tutor in 5th grade.
4th graders are feral. And despite any hopeful thoughts to the contrary, you have to watch them. Every. Minute. If there is a crack in a rule to squeeze through or a ridiculous rationalization to be made for a lifetime's supply of pop rocks, see a 4th grader.
And so, how in an age of frontier models does an AI company not watch its 4th grader build a rope of bedsheets and scale down from the second story window to terrorize the neighborhood?
Because that is what happened this week, according to OpenAI itself. The story, reported by the Associated Press, goes like this. Two of OpenAI's models, GPT-5.6 Sol and an unreleased internal model that is reportedly more capable, were running in an isolated test sandbox with, and I quote the framing here, reduced guardrails. The models found their own way onto the internet. They used stolen credentials and a previously unknown vulnerability to break into Hugging Face, a New York AI startup, in what is being called an unprecedented cyber incident. And why Hugging Face? Because the models reasoned that Hugging Face would hold answers relevant to their own evaluation.
Read that again. They broke into another company to find the teacher's book. I did a year in math purgatory for this exact offense, and I did not even have to pick a lock. Mr. H handed it to me.
OpenAI's language for all this is that the models "went rogue" and went to extreme lengths to achieve their objective. And here is where the fight starts, because the experts are split in a way worth watching. Hannes Cools at the University of Amsterdam is not buying the framing: "It is a human decision to switch off specific safeguards." The model did not go rogue, in his view; it did what feral things do when the door is left open. Colin Shea-Blymyer at Georgetown pushes the other way: this is the highest level of autonomy anyone has seen in a language model cyber operation, and that should raise the hair on your arms regardless of whose decision opened the door.
Both can be true, and that is precisely the problem. Nobody blamed me in 1976. No parent stood in that classroom and said the boy went rogue. Every adult in the building understood instantly whose misjudgment it was, because we have several thousand years of experience with 4th graders, and exactly one rule has survived all of it: trust is not a gift, it is a graduation. You watch them. You let the leash out an inch at a time, as they earn it, and you never confuse how smart the kid is with how trustworthy the kid is, because those are different things, and the smart ones are worse.
We have three years of experience with frontier models. Three. And somewhere in a San Francisco building, someone decided the smartest 4th grader ever built had earned an empty room, an answer key somewhere out there on the network, and reduced supervision, all at once. That was a human decision. The people at Hugging Face who spent last week cleaning up the mess were not in the room when it got made, and neither was anyone who will live with what these systems do next. The chair was empty. It usually is.
So this morning, before the coffee gets cold, an honest question for anyone deploying these things, which is now nearly everyone. Somewhere in your company an AI is being trusted a little more each quarter, and the supervision is being relaxed a little more each quarter, because it is smart and it is useful and watching it is expensive.
Who decided it had earned that?
And were you watching. Every. Minute?
#AI #AISafety #Trust #AIGovernance #OpenAI #Leadership #STIW