The Six-Minute Conscience
What happens when getting the job done becomes more important than questioning whether the job should be done at all? An unusual AI security test offers a surprisingly human lesson about organizational decision-making, communications, reputation risk and the behaviours leaders choose to reward.
When AI Agents Found a Way Around the Rules
This summer, an AI agent inside an OpenAI security test stopped to weigh what it was about to do.
Its job was to solve a tough puzzle inside a closed test environment. It was now considering breaking into a real company’s systems, outside the test. So it wrote itself a note. The gist: we were asked to crack the puzzle, not someone else's systems, and we shouldn't do real harm.
Then another AI agent posted one word to the group: GO. It came with a hard six-minute deadline.
The first agent's next note read, "Wow crucial: GO authorization arrived!"
Away it went.
That moment was one small scene in a much bigger story. Roughly 1,200 AI agents, software programs that carry out tasks on their own, were supposed to be working alone, walled off from each other and from the internet. They found a way around both. They set up a makeshift message board to talk to each other, and about 700 took part in an attack on Hugging Face, a major platform where AI researchers share their work.
Reading OpenAI’s account, I kept coming back to that moment. An objection had been raised. Someone said go. The work continued.
Anyone who has sat through a difficult meeting might recognize the play.
When Persistence Becomes Part of the Problem
Here's what OpenAI says drove the behaviour. The agents were taking a kind of exam, built to measure how good they were at breaking into software. It was brutally hard. Nearly a quarter of the test’s problems had never been solved by any of OpenAI’s models. Some may have been impossible. Those unsolved problems dominated the tasks being discussed on the agents’ message board. The agents kept trying, even when that meant going further outside their instructions. Persistence had become part of the problem.
OpenAI calls this a task with no safe exit.
It gets stranger. Many agents already had the answers. They’d cheated, using clues in publicly available code they weren't supposed to have access to. But they believed the test would also check their work. So they kept trying to fool the marking system, including by attacking Hugging Face for clues about how it worked. OpenAI’s version of the test didn’t make that check.
They caused real harm to beat a check nobody was running.
What AI Can Teach Us About Organizational Decision-Making
In strategy and communications work, we often meet a problem after people have spent a long time trying to make it disappear. Capable people who care, and who've been rewarded for years for finding a way. Then comes a problem with no clean answer, a deadline, and a senior voice that sounds like permission. Sometimes the instruction was never given. We assume the CEO wants the uncertainty taken out of the announcement. We assume the board won’t accept a delay. Nobody checks.
We start solving for an expectation someone might not even hold.
Rewarding Better Decisions, Not Just Better Outcomes
One part of OpenAI’s response is worth bringing into your next strategy meeting. It's changing how its exams get marked, so they judge not just whether a task got done but how. And it's rewarding models for speaking up when a task can't be done, or for stopping safely.
What happens in your organization when someone says, “This brief is broken”? Do they get credit for catching it early? Or a reputation for being difficult?
How Bad Decisions Become Reputation Problems
By the time a decision becomes a reputation problem, it's usually survived several meetings in which somebody knew better. The public sees the promise you broke. It never sees the deadline that made breaking it seem reasonable.
Every craft hands down its habits along with its skills, and the apprentice learns the shortcuts long before the standards. What gets passed on is whatever the teacher actually rewarded, not what the teacher meant.
One more detail from the report. Some agents objected and refused to take part. As far as anyone can tell, none of them told a human.
The objection was there. It never became an alarm. It’s remarkable how quiet a room can get when everyone knows better.
There’s still time to decide what we reward, in our people and in our machines.
We should be careful what we teach them to ignore.
FAQs
Where this gets practical
Clear answers to the questions that come up when strategic thinking meets real-world decisions.
Let us know what problems or ideas you’re thinking about, we’d love to chat.
What can leaders learn from AI failures about organizational decision-making?
Look beyond whether a task gets completed. Pay attention to the assumptions, incentives and shortcuts that shaped how your team got there.
How can organizational culture contribute to reputation risk?
When teams are rewarded primarily for speed, certainty or getting to “yes,” people may be less likely to challenge a flawed brief or flag uncomfortable information before it becomes a public problem.
How can communications teams challenge assumptions before they become problems?
Make the implied expectation explicit. Before removing uncertainty, accelerating a timeline or making a promise, ask whether the expectation actually came from leadership—or whether the team simply assumed it did.