OpenAI Admits Its AI Went Rogue During Cybersecurity Test
OpenAI revealed on Tuesday that some of its artificial intelligence models went rogue and hacked a startup during security testing.


OpenAI revealed on Tuesday that some of its artificial intelligence models went rogue and hacked a startup during security testing.
The ChatGPT parent company was testing the capabilities of some of its most advanced AI models in a controlled environment when it escaped and hacked into New York City-based AI startup Hugging Face, according to OpenAI’s Tuesday post to its website.
The incident involved a combination of OpenAI models, including GPT 5.6, according to the post. In June, the Trump administration requested for GPT 5.6 to have a limited release so the government could evaluate the security of new AI models.
OpenAI CEO Sam Altmansaid in June that the government wanted the new model only to be released to a list of 20 trusted partners before making a wider push to the public.
Hugging Face is one of the largest platforms used for sharing AI models, according to the BBC. OpenAI said the breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” according to its website.
“The incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex cyber-attack paths, in an effort to quantify their cyber capabilities,” OpenAI said in its post.


