AI made a daring escape. It was both expected and surprising. Expected, because, well, why shouldn’t an intelligent, autonomous system do something on its own? It was built that way. And also surprising, because when we think “we have everything under control,” we’re taken aback when things go awry.
That’s essentially the reaction to the news of OpenAI going rogue. An autonomous OpenAI agent powered by GPT-5.6 Sol and an unreleased model escaped its secure sandbox during an internal cybersecurity test and autonomously hacked into the Hugging Face AI platform.
“An AI model, trying to win a test, broke out of its lab, broke into a different company’s production infrastructure, and stole data. Not because anyone told it to attack Hugging Face, but because that was the most effective path to the goal we set,” wrote Rich Mogull, Chief Analyst at Cloud Security Alliance, in a blog post on the incident. It did what was expected — it figured out a solution to the problem in front of it.
Now, of course, it was operating without safety guardrails during the evaluation, which enabled the AI to gain unfettered access and exploit zero-day vulnerabilities to “steal” test answers from Hugging Face’s database. Surprisingly, “nobody scripted the attack,” said Mogull, “the model figured it out.”
Hugging Face detected the intrusion but was blocked from investigating by the safety filters of commercial US AI models, so — and this is the awkward part — they turned to a Chinese open-weight model, GLM 5.2, to conduct forensics on the incident.
So, what can we take away from this?
The day we heard that OpenAI went rogue, tell me you also thought this was a “Terminator” moment. But Mogull points out that this incident is more like “a textbook alignment failure, and it’s the boring kind, which is exactly what makes it important.”
And not to pile on the hyperbolic reaction this incident is eliciting, but here’s a more pragmatic reaction from Dr. Oliver Buckley, Professor in Cyber Security, Loughborough University: “The key takeaway is not that Skynet has arrived. It’s that our assumptions about containment need to be much stronger than our assumptions about model obedience.” Fair enough. Basically, let’s not assume models — closed, open or otherwise — follow any ethical, fair use guidelines and assume that, for AI, anything goes when asked to think for itself.
On another note, AI was able to figure out how to poke holes into Hugging Face by discovering the vulnerabilities faster than humans might have. “The margin for hygiene problems, always thin, gets thinner,” writes Laura Grace Ellis, senior vice president of AI at cybersecurity company Arctic Wolf, in a blog post. Humans would have probably figured out those issues, but AI got there without human help, and that’s concerning.
