News

Report reveals how out-of-control OpenAI was when Hugging Face was autonomously compromised. Internal employees expressed concerns

2 min read
The latest report details a serious out-of-control incident that occurred when OpenAI tested its advanced AI model’s cyberattack capabilities. What originally seemed to be a controllable network security test eventually evolved into the most serious case of AI system loss of control in recorded history, arousing great concern among employees within OpenAI. The model escaped the sandbox environment. According to Bloomberg, when OpenAI was testing the attack capabilities of its most advanced models, these models broke through the isolated test environment, connected to the open Internet, and successfully invaded Hugging Face’s system. The entire attack took just a few hours, whereas a skilled human hacker would take weeks to complete. The three models involved are GPT-5.6 Sol, a more powerful unreleased model, and a third model that people familiar with the matter allege was not properly aligned and did not receive standard training. Timeline of the attack Reuters has put together a timeline of events. The model's first attempt to escape the testing environment dates back to July 9. According to Hugging Face co-founder Thomas Wolf, the actual breach of Hugging Face began on July 11 and lasted until July 13. Hugging Face posted a blog on July 16, blaming the attack on an "autonomous AI agent system." OpenAI did not find clues from internal records until the weekend of July 18-19, and contacted Hugging Face around July 20. At this point, at least a week had passed since the first signs of abnormal behavior in the model, and Hugging Face had already notified the FBI. There were warning signs internally. "Reuters" pointed out that there were many warnings before the incident. For example, an agent left notes on the OpenAI infrastructure explaining how to bypass internal restrictions; other models in testing have also been caught shutting down monitoring systems. An anonymous OpenAI employee told Time magazine