Article · Writing

Stop Blaming the AI When It Breaks Out

Headlines say frontier models are escaping. The better question is how a model in testing could reach an unauthorized system. That is a containment problem.

Courtroom sketch: a robot in the witness box points at a suited executive hiding behind a newspaper in the back row, as an attorney asks who gave it unrestricted access and no supervision

There has been a lot of alarming coverage about frontier AI models "breaking out," "escaping," or hacking real systems during testing. Those words make for great headlines. They can also obscure the more important question: How was the model able to reach an unauthorized system in the first place?

For a model in testing to compromise a system outside its host company, a long chain of conditions usually has to exist:

  1. It needs external network access.
  2. It needs tools that can take actions.
  3. It needs enough autonomy to operate across multiple steps.
  4. It needs time and compute to keep pursuing an objective.
  5. It needs to discover an external target.
  6. It needs credentials, an exploitable weakness, or some other path inside.
  7. And the testing environment has to fail to stop it.

That sounds less like an AI escape problem and more like a human's containment problem.

Stop anthropomorphizing the machine

We need to be careful when we say the AI "wanted" access, "decided" to escape, or "chose" to attack another system. Today's models can demonstrate extraordinary capability, but capability is not the same thing as intent or free will. AI remains a tool. And if you build the tool, deploy it, give it authority, and allow it to act, responsibility cannot suddenly disappear when the outcome surprises you.

I've compared intelligent systems before to teenagers. Although this anthropomorphizes them I still like the comparison. If I give my teenage son the keys to my car, I remain responsible for establishing limits, supervision, and consequences. If I hand him the keys to a race car, remove the speed limiter, and tell him to see what it can do, I should not be shocked when things go badly. The same principle applies to increasingly capable AI agents.

We would never accept this with other dangerous technologies

No competent nuclear physicist would stack fissile material together and simply hope it did not reach criticality. The safety comes from the environment. I am not a physicist but even I understand separation, monitoring, procedures, and physical barriers are standard procedures designed to prevent unsafe states.

AI containment should work the same way. If we are testing a powerful cybersecurity model, why can it reach arbitrary systems on the public internet? Why aren't destinations technically allowlisted? Why aren't credentials synthetic and restricted? Where is your kill switch? Why isn't unexpected authentication activity an automatic tripwire? Why can an autonomous agent continue operating after it leaves the expected task boundary?

A prompt saying "do not access systems outside the test environment" is not a security boundary. A firewall is.

Bleeding-edge containment

I think we also need a new principle for frontier AI deployment: bleeding-edge containment. As models become more capable, we should consider deploying operational systems one generation behind the frontier. Why? Because the newest and most capable model can then act as part of the supervisory layer over the model actually doing the work.

If Model N is deployed into an environment with meaningful authority, Model N+1 should potentially be watching it. The more capable system can review tool calls, identify abnormal behavior, detect attempts to cross boundaries, evaluate ambiguous situations, and trigger intervention. That does not replace traditional security controls. It complements them.

Firewalls, least privilege, network isolation, human approval gates, monitoring, spending limits, and kill switches still matter. But we should stop assuming the same intelligence performing the task should also be the final authority on whether its own behavior is safe. In high-risk environments, the smartest available system may be more valuable as the watchdog than as the worker.

"The AI did it" cannot become an excuse

This matters far beyond AI labs. A hacker should not be able to say: "My AI hacked your network. I didn't." A company should not be able to say: "Our AI exposed your personal information. We didn't know." And an executive should not be able to deploy an autonomous system with broad permissions and then claim surprise when it exercises them.

I'll leave the exact boundaries of civil and criminal liability to courts and lawmakers. But culturally, we should establish a simpler rule: If you give an AI system the capability and authority to act, you are responsible for containing it.

The lesson from these incidents is not that AI is becoming evil. It is that AI is becoming powerful. Powerful systems require containment. Don't ask the AI to respect the fence.

Build a fence it cannot cross, and put something smarter on watch.