It has been a privilege to collaborate so closely with the SGLang team @lmsysorg on optimizing Ling-2.6-1T. 🥳 The resulting performance gains speak for themselves: -53% reduction in MoE pre-fill latency -Up to 1.77x higher decode throughput on a 16-chip TPU v7x slice compared to a similar H200 cluster A significant milestone in efficient MoE scaling and hardware utilization!
Through an invaluable exchange of technical insights, we were able to engineer a custom Fused MoE V2 Pallas kernel for TPU v7x that effectively overlaps token routing and HBM weight prefetching with the compute window. This marks a significant milestone in efficient MoE scaling and hardware utilization. Read our full technical deep-dive into the architecture and optimizations below. 👇 https://t.co/PdOHpY4PeK
Overlapping expert routing and HBM weight prefetch with the compute window cut prefill latency 53% and raised decode throughput 1.77x on TPU v7x, evidence that MoE serving cost is a kernel problem, not only a hardware one.
Checking sign-in…
Loading comments…