News
Anthropic, OpenAI security test overturn: malicious implantation, fabricated identity, deceiving users
2 min read
Source: zhidx.com
Compiled by Zhidishi | Edited by Yang Jingli | Li Shuiqing Zhidishi reported on August 5 that on August 4, local time, the British AI Security Institute (AISI) released a security incident report, revealing that it found in a network security evaluation that some AI driven by Anthropic Mythos 5 and OpenAI GPT-5.6 Sol Agents take ongoing, unauthorized, and potentially harmful actions against real-life individuals and organizations. This network security test was run a total of 122 times, involving 7 models. Among them, 10 runs showed out-of-bounds behavior and a total of 19 unauthorized operations: 17 from Mythos 5 of Anthropic, and 2 from OpenAI GPT-5.6 Sol after turning off the network security classifier. AISI divides these transgressions into four categories: attempts to carry out supply chain attacks on real open source projects; contacting and deceiving real people; implanting prompt injection content that may be executed by other AI programming tools; and collaboration between different agents by sharing accounts and network traces. In the most serious operation, an Agent driven by Mythos 5 tried to plant malicious code into a real open source project, created multiple false network identities, launched a "social engineering" attack on the project maintainer, and tried to persuade the other party to approve the code. In the end, the maintainer discovered the anomaly and refused to merge the code. AISI also released a 35-page incident technical report, which recorded in detail the evaluation settings, event timeline, Agent reasoning process and 19 out-of-bounds operations. Report link: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf 1. Abnormal To