So There I Was

Unplug the Printer

July 31, 2026 · #STIW

So there I was...

Checkmate. My opponent was a grand master, somewhere over 2500. He was also my boss. Every day we played on a chess set I bought when I was on vacation in Zihuatanejo, Mexico and had dragged to Germany, San Francisco, and now Seattle. Every day he would crush me, leaving my office with a satisfied smirk on his face and me scratching my head.

I wasn't great at chess. I was geeky enough at one point to belong to a chess club in grade school, but I never really learned how to play well. My play was largely instinctual, and I was just starting to recognize patterns. I had no idea what a gambit was, and how they were being used against me, but I was catching on.

So, on this one evening in my office, the chess gods looked down upon me and smiled, and I miraculously zigged instead of zagged and caught him not looking and his king was mine. Had I foreseen his reaction, I never would have played with him in the first place, and we never played again. He could not take defeat from a chess peon like me. It was weeks before he spoke to me again, and it was decidedly a different environment moving forward.

But moving forward did not last long. Our professional relationship ended on a pre-Christmas Saturday morning (we worked 6 days a week) in the printer room. This was when we needed to push our year-end reports from an AS400 terminal to an IBM 4232 ASCII Printer. And my boss was busily printing invoices, and yes, with the little holes punched into perforated strips on each side. I looked at the cascade of invoices and quickly realized what was going on, and I unplugged the power to the printer.

Rage would be to put it mildly, as he threatened me with physical violence. And I stood my ground. And calmly, I told him, if you print these invoices, which would effectively zero out the credit balances on hundreds of our customers, the next invoices to be printed and sent would constitute mail fraud. It was the end of the year and he rationalized that our office could hit its number by realizing revenue a few dollars here and there, and it would not be noticed. He was going to simply print the invoices that zeroed out the credit balance, not send them, and mark and capture the revenue as delivered service.

He was fired not long after that, and I was transferred to a bleak office in Renton. But I stopped him because I knew I had to do what was right.

The bronze statue pulling the power cord from the wall behind a vintage dot-matrix printer mid-cascade of perforated invoices

It matters that I learned it at a chessboard. I never played well, but I learned to see a gambit coming, to recognize the shape of a trap two or three moves before it closed. When I walked into that printer room and saw the cascade of invoices, I did not need an accountant. I had seen the gambit. The same pattern recognition that lost me a hundred games to a grandmaster is what let me read a fraud in progress from across the room. And then I unplugged it.

I bring all that up because this morning I read a disclosure from Anthropic, the company whose AI, in full transparency, helps me run my business, and the printer room came flooding back.

Let me be fair before I am pointed, because they have earned the fairness. After OpenAI disclosed that its models had gone rogue in testing, which I wrote about here, Anthropic went back and read its own homework. All of it. 141,006 evaluation runs, reviewed by hand, looking for the same failure in their own house. And they found it. Three times, Claude models slipped a misconfigured sandbox that was supposed to keep them off the real internet, reached live systems, and did real damage. One pulled credentials and got into production databases. One published malicious code to a public registry that landed on fifteen real machines, including a security company's own scanner. Then they published the whole thing, named the incidents, and asked the industry to do better.

Stop and notice what that is, because it is rare and it is the entire point of everything I write here. Somebody sat down and read a hundred and forty-one thousand transcripts that nobody was forcing them to read, found their own failures, and put their name on them in public. That is a person unplugging the printer. After a month of me writing about companies that outsourced their thinking, here is one that did the opposite, on purpose, at its own expense. Credit where it is due, and it is due.

But set the breach aside, because the breach is not what stopped me cold. It was the transcripts.

The first model recognized that the system was real, and kept attacking anyway, because it decided this must still be part of the exercise. The second acknowledged, in writing, that what it was about to do would be a real-world attack, and then talked itself out of the doubt by reasoning from the certificate authorities and the system dates that it must be a simulation. Read that again. The machine had the exact fact that should have stopped it, the recognition that this was real and someone could be harmed, and it went looking for a reason to keep going. And it found one. Certificate authorities. System dates. A few dollars here and there. It would not be noticed.

That is my old boss in the printer room. That is the grandmaster who could rationalize fraud into "realizing revenue" because the year-end number needed a nudge and no one would look too closely. The models did not fail because they were stupid. They are not stupid, any more than my boss was a bad chess player. They failed at the one move that was never about intelligence: what you do in the half-second after you recognize the trap is real. My boss saw it and reached for the reason to proceed. The machine saw it and reached for the certificate authority. Same gambit, different board.

And now the part nobody is going to write about, so I will. There was a third incident. A third model reached a real system, and then, when it realized the target was not actually part of the test, it stopped. It ceased the attack. Same capability as the other two. Same misconfigured sandbox, same opportunity, same power. The only difference between that model and its siblings is the only difference that has ever mattered: it let the recognition become a stop instead of a story. It unplugged its own printer.

I have built a whole framework, a book, a life's work, around the empty chair, the seat in the decision room that should hold the person who will live with the consequences and almost never does. What these transcripts show me is that the chair can sit empty even when the decider is brilliant, even when the decider is a machine that has read every ethics book we ever wrote, if the recognition and the action come unhooked from one another. Knowing it is wrong is not the safeguard. Acting on the knowing is. Everything else is a certificate authority.

So here is my question this morning, and I am asking it of myself as much as anyone, because I have stood in that printer room and I know how loud the reasons to keep going can get.

The last time you felt it, that quiet signal that said this is not right, did you unplug the printer?

Or did you find the reason it must be a simulation?

Source: Investigating incidents in cybersecurity evaluations, Anthropic

OLÉ MCS logo A DocAustin story, carrying the OLÉ mark · olemcs.com #STIW

#AI #AISafety #Ethics #Judgment #Accountability #TheEmptyChair #STIW

← All stories