News

DeepSeek V4 Pro won the second place in the programming evaluation: the score was only lost to Kimi K3, but the cost was only one-fourth of the opponent. The latest SuperCLUE-Terminal Chinese intelligent terminal programming evaluation list released by the CLUE team showed that DeepSeek-V4-Pro-0813 won

3 min read
DeepSeek V4 Pro won the second place in the programming test: the score was only lost to Kimi K3, but the cost was only one-fourth of its opponent. The latest SuperCLUE-Terminal Chinese intelligent terminal programming evaluation list released by the CLUE team showed that DeepSeek-V4-Pro-0813 scored 51.52 points, ranking second among all participating models, only behind the top-ranked Kimi K3. What's more noteworthy is that it achieves this result while maintaining extremely low calling costs, and its cost-effectiveness advantage is particularly prominent among high-scoring models. This list is specially designed for domestic developers, all using local real programming tasks, and uniformly using Claude Code as the operating framework, to fairly compare the complete capabilities of major models in understanding Chinese requirements, planning tasks, calling tools, and troubleshooting and modifying codes. The realism of the evaluation comes from its task setting: each test question requires the model to complete 50 to 110 rounds of interactions, and a complete run of a single question takes half an hour to an hour and a half, which can highly restore the complex workflow in daily development. In this round of the list, Kimi K3 ranks first with 60.61 points, followed by DeepSeek-V4-Pro-0813, which has a score higher than GLM-5.2’s 48.48 points and the old version DeepSeek-V4-Flash’s 46.46 points, firmly ranking in the first echelon. The cost of a single question is only about 1.42 yuan, which is about one-quarter of the Kimi K3 and one-sixteenth of the GLM-5.2. At a time when head model competition scores are often only a few points apart, this order of magnitude cost gap means that small and medium-sized teams can afford "near-top" code capabilities without having to pay high computing power for each interaction. Judging from actual measurement scenarios, the model covers real needs such as front-end pages, 3D games and engineering script modifications, and can complete closed-loop game development logic. Collision, scoring, and settlement functions all operate normally. Page color matching and interactive feedback also get rid of the templated AI style. A few creative drawing questions will have minor structural flaws, but the overall completion level is still at a leading level. For budget-sensitive small and medium-sized enterprises and independent developers, DeepSeek-V4-Pro-0813 redefines the cost-effectiveness benchmark of programming models with "second-place scores and one-quarter price" - when top capabilities no longer mean top expenses, the threshold for the popularity of AI programming is truly being lowered. via AI News (autho