News

Actual measurement DeepSeek-V4-Pro: stronger than logic closed loop, weaker than vision and interaction

1 min read
Source: aixq.cc
On the card table of large models, DeepSeek has always been a very special existence: the price is less than a fraction of the US closed-source flagship, and relying on the ultimate sparse architecture, the cost of using cutting-edge capabilities has been pushed to the floor. Faced with real engineering tasks with complex rules and clear acceptance criteria, can it directly deliver a result that can be run, interacted, and played? This time, we let DeepSeek-V4-Pro stand on the same stage as Kimi K3 and Qwen3.8-Max and run the same five hard-core scene questions. Five real engineering task tests To judge whether a model is capable of complex tasks, at least three levels must be examined: whether it truly understands the rules; whether it has established the correct data and status system instead of using static pages or pre-set