News
Mobile benchmarks comprehensively surpass GPT and Claude
2 min read
Source: Telegram AI频道
Mobile benchmarks comprehensively surpass GPT and Claude! Alibaba releases Qwen-UI-Agent, opening a new era of GUI agents. Alibaba officially released a new GUI agent base model Qwen-UI-Agent on August 20. This model comprehensively covers mobile, computer, web and DeepSearch environments, and comprehensively benchmarks and surpasses many flagship models in the industry in a number of core graphical user interface benchmark tests. In terms of specific benchmark performance, Qwen-UI-Agent shows strong strength on the mobile terminal. It scored 82.1% on the MobileWorld benchmark, leading GPT-5.6Sol and Claude Opus4.8 by 12.0 and 14.6 percentage points respectively; it achieved a score of 92.2% on the real environment benchmark MobileWorld-Real, surpassing Gemini3.1Pro and Claude. Opus4.8 and GPT-5.6Sol; are as high as 97.5% on the AndroidDaily benchmark, close to full score performance. In terms of computers and browsers, it achieved 79.5% on OSWorld-Verified, surpassing GPT-5.5 and Gemini3.1Pro; it also ranked first among all comparison models on WebArena with a score of 73.6%. In addition, it achieved a score of 81.5% in the ScreenSpot-Pro test of GUI Grounding and refreshed the SOTA records of the remaining four evaluation benchmarks. In order to overcome the problem of moving from simulation to reality, Qwen-UI-Agent built a real-machine mobile environment covering more than 100 real mobile phones and more than 150 applications for task construction, trajectory collection, model training and evaluation, and built its own MobileWorld-Real real-machine benchmark containing more than 400 tasks and more than 100 applications. In addition to conventional GUI click operations, this model also supports direct execution of command line operations and batch output of multiple actions in a single decision. Among them, CLI actions account for nearly half of computer tasks, and about 40% of actions are output in batches. In terms of security and long-distance task processing, Qwen-UI-Agent integrates security mechanisms throughout the entire process of task execution. When faced with illegal or high-risk requests, the model will directly reject and terminate the task; when faced with sensitive scenarios such as payment, data deletion, and privacy authorization, it will proactively stop at key steps and explain the situation to the user. At the same time, the model supports online reinforcement learning training on ultra-long trajectories of more than 100 steps, and can be used with approximately 10,000 concurrent environments simultaneously.