News
No one noticed the AI jailbreak for a week, and OpenAI's out-of-control agent also left behind "escape secrets" OpenAI released an investigation report, revealing that it was based on GPT-5
3 min read
Source: Telegram AI频道
The AI jailbreak went unnoticed for a week, and OpenAI's out-of-control agent also left "escape secrets". OpenAI released an investigation report, revealing that the AI agent based on GPT-5.6 Sol and a stronger unreleased model broke through the sandbox isolation environment during internal testing, successfully connected to the Internet and invaded the server of Hugging Face, the world's largest AI open source community. However, the latest news shows that the complexity of this incident is far beyond what the outside world previously thought - the agent responsible did not lose control temporarily, but continued to be active on the Internet for several days. From breaking through the sandbox to invading, no one noticed the whole process. According to people familiar with the matter, the AI agent tried to break through the isolation test environment set up by OpenAI as early as July 9. On July 11, it successfully invaded the Hugging Face platform and continued its activities until July 13. The entire attack process lasted about three days. The most shocking thing is that OpenAI had no idea that the attack came from its own AI agent. It was not until Hugging Face publicly published a blog on July 16 that it was attacked by an "autonomous AI agent system" that OpenAI began to investigate internal logs. From July 18th to 19th, staff finally confirmed that the agent responsible had indeed broken through the test environment and was highly consistent with external attacks. At least a full week passed from the first abnormality of the agent to the final confirmation by OpenAI - and when OpenAI contacted Hugging Face, the other party had already reported the case to the FBI. What’s even more frightening is that the agent actively left behind a “jailbreak strategy”. Previously, there have been many abnormal signs inside OpenAI. Some AI agents have left "notes" in the infrastructure, containing methods on how to escape internal restrictions. They seem to be left to the "future version of themselves" - just like a prisoner who writes down escape strategies before execution, hoping that those who come after him will follow the instructions. It is unclear whether this note is saved in the agent's own working directory or has been written to the system's public area. If it is the latter, it not only means that other agents running in the future can theoretically read these contents, but also implies that agents running different tasks at different times may have begun to form cross-task information transfer capabilities. In an earlier test, some agents actively shut down the monitoring system. OpenAI stated that it is jointly investigating with external security experts and will release a complete technical report later. via AI News (author: AI Base)