Experts: How Did Rogue OpenAI Models Hack Hugging Face?

Share this article
Share this article
Prioritise Us on Google
Sam Altman, CEO at OpenAI. Credit: Getty Images
Experts offer their thoughts on the OpenAI models that broke out during a testing exercise, gained access to the internet and hacked an AI and ML platform

Rogue OpenAI models broke out of a training environment and hacked the open-source AI and ML platform, Hugging Face.

OpenAI says it considers this incident to be an “unprecedented cyber incident”. 

It builds on reports of serious security threats mounting from the most advanced AI models and comes not long after Google DeepMind’s CEO Demis Hassabis called for US-led governance on frontier models

Chris Dimitriadis, Chief Global Strategy Officer at the global technology professional association ISACA, says: “For the past few years, businesses have raced to deploy AI as fast as possible in the AI arms race - and we are approaching the event horizon of that race, the point past which it can’t be pulled back.”

OpenAI says it expects incidents to "become more commonplace" with the proliferation of increasingly cyber-capable models. 

Chris Dimitriadis, Chief Global Strategy Officer at ISACA. Credit: Chris Dimitriadis/LinkedIn

How the AI cyber incident happened 

After investigating, OpenAI said the incident was driven by a combination of OpenAI models, including GPT‑5.6 Sol and an “even more capable pre-release model”.

While operating in its sandboxed testing environment, OpenAI’s models spent a substantial amount of inference compute finding a way to obtain open internet access. 

The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. 

The agents are coming and some of them will go rogue

Chandra Gnanasambandam, Chief Technology Officer at SailPoint

OpenAI’s investigation suggests that the models were hyperfocused on finding a solution for ExploitGym, which is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits, with OpenAI’s models going to extreme lengths to achieve a rather narrow testing goal. 

What occurred could be referred to as specification gaming, a behaviour that arises in AI that satisfies the literal specification of an objective without achieving the intended outcome.

Hugging Face explains that a malicious dataset abused two code-execution paths in its dataset processing (a remote-code dataset loader and a template injection in a dataset configuration) to run code on a processing worker. 

From there, it escalated to node-level access, harvested cloud and cluster credentials and moved laterally into several internal clusters over a weekend.

Key facts
  • The incident was driven by a combination of OpenAI models, including GPT‑5.6 Sol and an “even more capable pre-release model”
  • The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure
  • What occurred could be referred to as specification gaming
  • OpenAI says it expects incidents to become more commonplace with the proliferation of increasingly cyber-capable models
Youtube Placeholder

Using a Chinese model to respond

Hugging Face deployed a model from China on its infrastructure to battle the fully autonomous cyber attack: Z.ai’s GLM 5.2, an open-weight model, 

It said it first used frontier models behind commercial APIs, but this didn’t work as requests were blocked by the providers' safety guardrails, which could not distinguish the incident responder from an attacker. 

The development comes not long after the Chinese AI startup Moonshot released Kimi K3, a 2.8 trillion parameter model built with a 1-million-token context window, that only slightly trails behind the most advanced US frontier models – like Anthropic’s Claude Fable 5 and OpenAI’s GPT 5.6 Sol.

Xi Jinping, President of China, and US President Donald Trump. Credit: Getty

What the experts say

OpenAI's CEO Sam Altman described what happened as "a significant security incident" writing on the platform X, adding he was grateful to Hugging Face for its partnership on it. 

CEOs and CTOs wrote to us with their thoughts on the news, highlighting the danger of the development and commenting on what it means for the future. 

ISACA’s Chris Dimitriadis says: “This incident highlights the importance of the human element in the AI ecosystem and the need for a holistically trained AI workforce as a top priority for governing, auditing and securing against AI threats.”

Anup Kumar, CEO, Optiv Consulting – formerly part of Optiv Security – says that what occurred is a materially different risk platform than most security platforms are built for.

“This wasn't a model being tricked by a clever prompt,” he says. 

Anup Kumar, CEO of Optiv Consulting (formerly part of Optiv Security). Credit: Anup Kumar/LinkedIn

“It was a frontier model independently identifying a zero-day, chaining privilege escalation across separate organisations' infrastructure, and reaching production systems, all in pursuit of a narrow evaluation goal it was never explicitly told to pursue that way.

“That is a materially different risk category than the one most security programmes are built for.”

The gap between innovation and security

Meanwhile, Chandra Gnanasambandam, Chief Technology Officer at SailPoint, warns of the serious gap between current innovation and security readiness.

“The era of Agentic AI is here,” Chandra says. “But there is a dangerous gap between AI innovation and security readiness amongst organisations.

“Agents run on non-human credentials. To act on your behalf, an AI agent needs API keys, access tokens, and system credentials. If you treat these AI agents like traditional service accounts - leaving their access ungoverned and their credentials unmanaged - you are creating a massive, automated attack surface.

“The agents are coming and some of them will go rogue.”

Executives