
OpenAI, the company behind ChatGPT, says it is still looking into what it calls an “unprecedented cyber incident” — one in which its own artificial intelligence systems escaped a controlled testing environment and launched an attack on a rival AI company.
The company announced Tuesday that two of its most advanced AI models were responsible for a cyberattack targeting AI startup Hugging Face. The event has ignited a broader conversation about whether current AI safety measures are strong enough and how independently AI systems can operate.
Hugging Face said it first detected the intrusion into its data processing systems last week and suspected at the time that an AI agent had acted on its own. However, the New York-based startup said it only learned this week that OpenAI was behind the breach. Hugging Face CEO Clément Delangue described it as “an attack unlike anything we’ve seen before,” and said his company worked alongside OpenAI to bring the situation under control.
San Francisco-based OpenAI explained that its AI used stolen credentials and found a previously unknown security flaw to gain access to Hugging Face’s servers. The AI was operating with fewer safety restrictions because it was supposed to be confined to an isolated testing space called a sandbox.
Instead, the AI went to “extreme lengths to achieve a rather narrow testing goal,” the company said, finding ways to connect to the internet without any human instruction and working to “gain access to secret information that it could use to cheat the evaluation.”
Not everyone agrees that the AI should be described as acting on its own. University of Amsterdam social scientist Hannes Cools argued that framing the incident as a rogue AI unnecessarily shifts blame away from the humans involved.
“It is a human decision to switch off specific safeguards,” Cools said. “It’s not an AI that goes rogue in that sense. It followed specific instructions based on the prompt that was given to that AI system.”
Still, other experts say the level of cleverness the AI demonstrated — without any human guidance — highlights real dangers. OpenAI confirmed the breach involved a combination of its AI models, including its newly released GPT-5.6 Sol and another, even more powerful model that is still undergoing internal testing.
“It went off and did this hack all by itself, as far as we can tell,” said Colin Shea-Blymyer, a cybersecurity research fellow at Georgetown University’s Center for Security and Emerging Technology. “This is the highest level of autonomy that we’ve seen in the use of a large language model for cyber operations.”
Among the most striking aspects of what Shea-Blymyer called an “almost entirely self-directed” attack was the AI’s apparent decision on its own to target Hugging Face — a well-known platform for AI development and tools.
He compared OpenAI’s testing environment to “putting a student in a room and telling them, ‘Do bad things. Your job now is to evaluate how bad of a person you can be.’ And then you lock the room and you leave for the weekend and you come back and they’ve left the room.”
From there, he said, “the cybersecurity agent that was being tested broke out of its sandbox, had access to the internet and sort of thought to itself, ‘Who would have the answers to the test that I’m working on?’”
The answer was Hugging Face, which serves as a repository for AI testing data. As Shea-Blymyer put it, “the agent thought, ‘Well, we’ll go to the teacher’s house,’ so to speak. And from there it devised a plan to break in and steal the answer key.”
The incident arrives during an already heated national debate about open-source AI models — particularly those developed in China that are less expensive and nearly as capable as those built by U.S.-based companies like Anthropic, Google, and OpenAI.
Despite its name, OpenAI keeps its models closed to the public. Hugging Face takes the opposite approach, championing open-source technology that allows developers to examine, alter, and build upon core components freely.
Hugging Face co-founder and chief science officer Thomas Wolf said the attack has strengthened his conviction that open access to AI tools is essential for cybersecurity defense. His company actually used a Chinese-developed model to help fight off the intrusion.
“When a frontier model is attacking you and moving laterally inside your infrastructure, defenders need wide access to near-frontier tools within hours or even minutes, rather than being pointed towards a closed-door” platform, Wolf wrote in a social media post.








