News

Ascend 950 super node real machine debuts and wins awards: Huawei achieves the industry’s largest 1024 card scale

2 min read
Kuai Technology reported on July 18 that the Ascend 950 super node real machine, which had just made its public debut at the 2026 World Artificial Intelligence Conference, won the industry's major award immediately upon its debut, which shows the high attention and strategic weight of this product at this conference. Huawei China officially announced that the Ascend 950 super node (Atlas 950 SuperPoD) successfully won the SAIL Award, the highest honor at the 2026 World Artificial Intelligence Conference (WAIC). The SAIL Award, which stands for Outstanding Artificial Intelligence Leader Award, is the highest-level honor set by the World Artificial Intelligence Conference. It is also one of the most authoritative and internationally influential iconic industry awards in the global artificial intelligence field. All award-winning projects represent the top breakthrough technological achievements in the global AI field that year. At this year's WAIC exhibition, Huawei exhibited the fully configured Ascend 950 super node offline for the first time. This product relies on the self-developed Lingqu interconnection protocol and the newly upgraded super node architecture to launch the largest 1024-card computing power cluster in the entire industry, directly raising the upper limit of cluster support for domestic AI computing power to a new level. It can output surging computing power of 1 EFLOPS FP8 and 2 EFLOPS FP4, and has a 256TB global unified memory addressing space, which is fully sufficient to support the full-process efficient training of ultra-large billion-level models without the need for additional splitting and adaptation of the cluster architecture. The entire cluster relies on TB-level NPU interconnection ultra-large bandwidth and 3μs ultra-low RTT latency, which directly breaks through the communication bottleneck of various resources such as internal computing and storage in traditional large clusters. Through systematic innovation and optimization of the full-stack architecture, unified and efficient scheduling of global computing resources is achieved. Ultimately, the actual operating efficiency of the entire ultra-large-scale cluster is much higher than that of conventional computing clusters using traditional interconnection solutions at the same scale.