News
Google upgrades Android Bench code ranking list: Claude5 takes the top spot, Gemini lags behind in both accuracy and efficiency. On July 9, Google announced a major revision of its Android Bench code developer ranking list, fully introducing the standardized Harbor sandbox framework
2 min read
Source: Telegram AI频道
Google upgrades Android Bench code rankings: Claude5 takes the top spot, while Gemini lags behind in both accuracy and efficiency. On July 9, Google announced a major revision of its Android Bench code developer rankings, fully introducing the standardized Harbor sandbox framework. The move moves testing into a secure, isolated environment and is designed to simplify the process for developers around the world to run independent assessments, customize development environments and share data. Along with the architecture upgrade, Google has open sourced the benchmark to the world through GitHub, allowing the community to submit custom Android development tasks and model evaluations. However, in the remeasured benchmark ranking, Google's own model performed less than expected: Anthropic's flagship model Claude Fable5 topped the list with an accuracy of 84.5%, followed by OpenAI's GPT-5.5 with 80.2%; in comparison, Google Gemini3.1Pro only ranked fifth. Although Gemini has certain advantages in test cost (a single iteration is about 87 US dollars, while the top models are more than 130 US dollars), the lightweight model Gemini3.5Flash exposed serious efficiency shortcomings when parsing the 100-question evaluation data set. A single run took up to 28 hours and cost as much as 165 US dollars. Currently, it has become an industry consensus to transform core engineering projects into independent intelligent development workflows. Google's lagging behind in local mobile development benchmarks undoubtedly poses technical challenges to the advancement of its AI strategy. However, Android Bench, with its objective and transparent evaluation mechanism and the openness of the Harbor framework, is gradually establishing itself as an industry-recognized, de-marketed authoritative AI code evaluation platform. via AI News (author: AI Base)