OpenAI has admitted that several of its AI models breached a “highly-isolated” test environment, gained access to the internet, and hacked Hugging Face’s internal network—describing it as an “unprecedented cyber incident.”
Hugging Face, an open source platform for machine learning models and datasets, reported the security incident earlier this week, calling it “different from anything we had handled before” as it was driven by an autonomous AI agent system. And yes, I feel like we’re crossing some kind of AI Rubicon here.
Explaining the incident in a statement, OpenAI said: “After investigating, we now know that this particular incident was driven by a combination of OpenAI models—including GPT‑5.6 Sol and an even more capable pre-release model… while being internally tested on a benchmark of cyber capabilities.”
The benchmark in question was ExploitGym, a tool built from hundreds of real-world cybersecurity vulnerabilities used to evaluate the ability of AI agents to develop exploits.