Huawei showcased the Atlas 950 SuperPoD servers for the first time in WAIC after first announcing it at MWC in March. An Atlas 950 SuperPoD consists of 1,024 NPUs, which consists of 16x Atlas 950 racks with each rack having 64 Ascend 950 NPUs. Each 950 SuperPoD therefore boasts 1 EFLOPS FP8 and 98TB of HiZQ memory (Huawei’s custom HBM3E equivalent). An additional 158TB of second tier DRAM can be added to provide a total 256 TB of usable memory. The current Atlas 950 SuperPoD is designed for the Ascend 950DT, which is the version that focuses on decode and training that uses Huawei’s HBM-class HiZQ memory. The Ascend 950PR version is prefill focused as it uses HiBL memory, which is effectively LPDDR on package and therefore lower bandwidth. (1/4)🧵

Seven ports are used to connect other NPUs on the same tray over flyover cables. Seven ports connect across trays in the intra-rack 2D mesh over the backplane allowing each NPU0 to connect to all the other corresponding NPU0 within the rack. This leaves the remaining 4 ports to connect to a network of low-radix-switches in the backplane for further intra-rack connectivity. (2/4)

For inter-rack connectivity to the 1,024 NPUs SuperPoD, each of the 32 LRS has 4x100G going to 6 other corresponding LRS to 6 neighboring racks. The inter-rack connection uses Hauwei’s own StarMatrix pluggables 800G LPO pluggable transceivers. As they are LPO, optical DSPs are not required, which significantly reduces transceiver cost and power. On the roadmap, Huawei is still undecided whether to implement NPO or stick with 1.6T LPO pluggables for the next gen Ascend 960 system. (3/4)

For more details, please subscribe to our Accelerator & HBM Model https://t.co/iNX6TOZ3hd as well as AI Networking Model. (4/4) https://t.co/kqt0IiSiFi
Huawei's 1,024-NPU Atlas 950 SuperPoD reaches 1 EFLOPS FP8 with 98TB of HBM-class memory, and splits prefill onto LPDDR-class parts and decode onto HBM-class ones. Its LPO optics skip DSPs, cutting transceiver cost and power at rack scale.
postHow NVIDIA's LPU splits prefill and decode, and what GPTOSS 2T actually measures
postAMD's inference gap is a software and tooling problem, not a silicon one
articleMost Neoclouds Suck At Security OpenAI vs HuggingFace, Container Escapes,…Checking sign-in…
Loading comments…