News
DeepSeek V4 Pro official version reversed: really strong, full health performance is restricted
3 min read
Source: news.mydrivers.com
Kuai Technology reported on August 15 that it has been three days since the official version of DeepSeek V4 Pro was released. There are still many controversies, especially the official test performance is directly catching up with Fable 5, but the actual test results of netizens are not that strong, although it is indeed improved compared to the preview version. Coupled with the huge price increase of DeepSeek's API, the cache hit price has increased by up to 11 times, and the overall cost of use has increased significantly. For a time, the reputation has seriously declined, and Liang Sheng's name has been lost. However, new research has appeared in the past two days, showing that the actual performance of the official version of DeepSeek V4 Pro is indeed very strong. It can not only reach the level stated by the official, but can even reproduce the amazing ability in the grayscale test in July. It can be said that it has returned with full blood. The reason why the performance of the official version of DeepSeek V4 Pro is very different is that it has an important relationship with its thinking chain. In online testing, several situations such as we need and let me appeared, and the performance of different thinking chains fluctuated greatly. After verification by the open source community, it was stated that the official version of V4 Pro relies heavily on the first round of tool targeting, so there is a solution. The open source community made a dsh-anchored-standard plug-in, which adopts the "first round anchoring + dynamic promotion" strategy - only 2 core tools are exposed in the first request to anchor high-quality inference tracks. Once the model initiates the first Tool Call, all 25 standard tools are immediately unlocked in subsequent rounds. After verification, in the Windows native Project2 test, this solution continuously ran high scores of 98 and 99, successfully entering the score band of cutting-edge models such as Fable. This experiment shows that the core capabilities of V4 Pro have not been lost, but there may be significant overfitting to the specific Harness scaffolding and tool exposure environment in the reinforcement learning (RL) post-training phase. At present, this plug-in is open source, and the address is here. For detailed instructions, you can also refer to the records of the open source community. The analysis inside is very comprehensive. If you are interested, you can learn more. In addition, the performance of the official version of V4 Pro is comparable to other