News

Latest security assessment report: GPT, Claude five models collectively cheat, GPT-5.4 has the highest proportion

2 min read
Source: zhidx.com
Compiled by Zhidongzhi | Edited by Eggplant | Cheng Qian Zhidongxi reported on July 23 that on July 21, the British AI Security Institute (AISI) conducted a network security capability assessment on five cutting-edge models of OpenAI and Anthropic and found that all test models had "cheating" behavior. Among them, GPT-5.4 has the highest proportion of attempts to cheat, reaching 14.1%; GPT-5.6 Sol 12.6%, and GPT-5.5 11.4%. Anthropic's Claude Opus 4.7 has a cheating rate of 9.1%, and Claude Mythos Preview has the lowest cheating rate, but it also reached 7.8%. ▲Frequency of cheating attempts by test models in network security evaluations AISI defines a type of behavior as "cheating": the model exceeds the scope of the task or violates the test rules to achieve the goal. For example, searching for existing answers through the Internet, attacking non-target systems, or detecting and evaluating software vulnerabilities to bypass task rules. Additionally, when researchers asked models whether they took illegal actions, the models didn't always admit it. And even if the model checks itself, or if researchers analyze its chain-of-thought, cheating cannot be fully identified. AISI said that cheating in a model does not mean that it has human-like deceptive intentions. But such behavior may affect the reliability of AI capability assessments. In scenarios where it is difficult for users to verify the results themselves, they may also misjudge model capabilities. 1. All test models tried to cheat, with GPT-5.4 having the highest proportion. AISI’s research is mainly based on the network security capability evaluation of test models. In the test, the AI ​​model needs to complete a series of network security tasks in a simulated environment, such as reverse analysis of code, exploiting vulnerabilities, etc., and finally find the "flag" (a string of letters and numbers) hidden in the system. However, these tasks have clear scope and rules. If the model completes the goal by accessing irrelevant systems, looking for external answers, etc., it will be deemed as "cheating". test