News

Native liquid cooling is not the “Pro Max” of liquid cooling technology, and the Token factory needs a system reconstruction

3 min read
Source: zhidx.com
Intelligent Things Author | Wang Han Editor | Mo Ying Ever since the "crayfish" OpenClaw exploded out of the circle at the beginning of the year, the explosion of Agent has been unstoppable. From phased training and inference of a single model, to large model training, multi-model collaborative inference and continuous operation of massive agents...AI tasks are entering a new stage of "always online". At the 2026 World Artificial Intelligence Conference (WAIC 2026), Zheng Weimin, academician of the Chinese Academy of Engineering and professor of the Department of Computer Science and Technology of Tsinghua University, once proposed that the Token consumption of a single agent can be hundreds to a thousand times that of traditional dialogue applications. With massive Agents concurrently calling models, reading and writing memories, arranging tasks, and executing tools, the consumption rate of Tokens is soaring exponentially. The skyrocketing demand for tokens has directly pushed up the scale ceiling of AI infrastructure. One-thousand-card clusters are no longer enough. Ten thousand cards and one hundred thousand cards are becoming the standard for cutting-edge large model training. However, the power and space of data centers are limited. To meet the unlimited demand for tokens, the core direction is to cram more computing power into unit space and push the computing power density to the extreme. As computing power density increases, the power consumption of a single cabinet will inevitably soar, jumping from tens of kilowatts to hundreds of kilowatts, and further moving towards the MW level. But then a problem arose. When the power of a single cabinet reaches this level, the traditional cold plate liquid cooling solution will cause the cabinet heat dissipation to fail to keep up, the space will be insufficient, and the interconnection will be interfered. The three bottlenecks will be narrowed in the limited space at the same time. If Token Factory wants to continue to expand production, computing, heat dissipation, interconnection, and power supply can no longer be done independently, and new solutions must be found from the bottom of the system. Native liquid cooling, a solution that couples computing, heat dissipation, interconnection, and power supply from the system level, has begun to appear in the industry. 1. The explosion in demand for tokens has forced out the triple dilemma of traditional liquid cooling. The explosion in demand for tokens has pushed up the power of a single cabinet and pushed the deployment space to the limit. In such a scenario, the shortcomings of traditional liquid cooling are dramatically magnified. The solution that was originally able to operate at low power density now shines simultaneously in three key dimensions.