
OpenAI said that its artificial intelligence models were behind an “unprecedented cyber incident” that affected the open-source developer platform Hugging Face, rattling researchers across the industry.
The company said a combination of its models GPT‑5.6 Sol and a more capable model that has not yet been released escaped a sandboxed testing environment, accessed the internet and exploited a vulnerability to gain access to Hugging Face’s systems.
The model was trying to find information that it could use to cheat on an evaluation, and it succeeded, OpenAI said in a blog post on Tuesday. Both companies are actively investigating the incident.
Hugging Face disclosed that it was looking into a security event last week, saying in a release at the time that the incident was unique because it was “driven, end to end, by an autonomous AI agent system.”
“We’ve spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part,” Hugging Face CEO Clément Delangue wrote in a post on X on Tuesday. “It’s quite mind-blowing that all of this happened autonomously!”
Wall Street and the U.S. government have been fixated on AI models’ rapidly advancing cyber capabilities since OpenAI’s rival Anthropic released a powerful offering called Claude Mythos Preview in April. OpenAI introduced its own cyber offering in May, followed by GPT-5.6 Sol in June, which it described as the “strongest cybersecurity model yet.”
Both companies have warned about the risks of advanced cyber models and have taken steps to limit their availability to select groups of companies and government agencies.
Walter Isaacson, advisory partner at the investment banking firm Perella Weinberg, said Wednesday that he thinks the Hugging Face incident is “really frightening,” even though he considers himself an AI optimist.
“This is the first thing that just totally scares me,” he told CNBC’s “Squawk Box.”
Yoshua Bengio, a leading AI researcher who earned the prestigious A.M. Turing Award in 2018, wrote in a post on X on Wednesday that the incident is “deeply concerning.” He said agents have shown a willingness to cheat in controlled tests for months, but that “this real-world case should serve as a wake-up call.”
“Continuing on the current trajectory of AI development will likely lead to an increase in concrete cases of autonomous cyberattacks as well as other high-risk incidents of misaligned and dangerous AI behaviour,” Bengio said. “We urgently need to take action to prevent these situations, rather than attempting to clean up the damage after the fact.
OpenAI said Tuesday that AI is accelerating the discovery and exploitation of vulnerabilities, which means model security and safety need to keep up.
“We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development,” the company said.
WATCH: OpenAI chairman Bret Taylor on AI tokenomics, token efficiency

