News

AI computing power competition enters the system era! ZTE OEX super node released, redefining AI factory base

2 min read
Source: zhidx.com
Intelligent things Author | Chen Junda Editor | Mo Ying In this year's WAIC, supernode is still the most popular keyword in the AI ​​infrastructure exhibition area. Along the way from the entrance of the exhibition hall, almost all domestic AI infrastructure manufacturers have placed their super-node products in the most conspicuous position. Why is everyone talking about "super nodes"? As trillion-parameter models continue to emerge, it has become increasingly difficult for the GPU alone to determine AI system performance. What really affects the efficiency of model training and inference begins to become how GPUs are connected, how servers are organized, how networks collaborate, how software is scheduled, and how the entire data center operates efficiently as a whole. The competition in AI infrastructure is also moving from single-point performance competition to system capability competition. It is against this industrial background that ZTE officially released the new generation OEX super node at the "Extreme Collaboration Shanghai Ecosystem" full-stack intelligent computing ecological forum held during WAIC, providing a solution for the system era. At the same time, the key technologies and applications of the domestic high-performance Matrix super node based on the OEX+dOCS architecture, jointly created by ZTE and partners such as Xizhi Technology, Biren Technology, Muxi Technology, Suiyuan Technology, Tianshu Zhixin, also won the 2026 WAIC SAIL Award. What problems does the OEX super node solve that have long plagued AI infrastructure? And why has it become one of the most watched products in the WAIC AI infrastructure field this year? 1. AI computing power competition has entered the system era, and traditional architectures are in urgent need of iteration. In the past few years, the development speed of large models has far exceeded the speed of chip performance improvement. The scale of model parameters is approaching trillions, training data and inference requests continue to grow, and factors such as power consumption, memory capacity, and interconnection bandwidth are also beginning to restrict each other. It is difficult to achieve linear computing power growth by simply stacking more GPUs. Whether it is token circulation in the MoE model or parallel strategies such as Tensor Parallel and Expert Parallel, the communication volume increases exponentially. If the interconnection efficiency is insufficient, no matter how powerful the GPU is, it can only wait for data