A fused MoE Pallas kernel cut Ling-2.6-1T prefill latency by half
- Source
- Ant Ling
- Date
It has been a privilege to collaborate so closely with the SGLang team @lmsysorg on optimizing Ling-2.6-1T. 🥳 The resulting performance gains speak for themselves: -53% reduction in MoE pre-fill latency -Up to 1.77x higher decode throughput on a 16-chip TPU v7x slice compared to a similar H200 cluster A significant milestone in efficient MoE scaling and hardware utilization!
Through an invaluable exchange of technical insights, we were able to engineer a custom Fused MoE V2 Pallas kernel for TPU v7x that effectively overlaps token routing and HBM weight prefetching with the compute window. This marks a significant milestone in efficient MoE scaling and hardware utilization. Read our full technical deep-dive into the architecture and optimizations below. 👇 https://t.co/PdOHpY4PeK
Context
Serving a model requires routing each to the right subset of experts and pulling those experts' weights from memory before computing on them; done sequentially, that routing and fetching step can stall the chip. Ant Ling reports that, working with the SGLang team, it built a custom Fused MoE V2 Pallas kernel for Google's TPU v7x chip that overlaps token routing and HBM weight prefetching with the compute window, for its Ling-2.6-1T model, whose technical report had shipped two days earlier. The company reports this cut MoE prefill latency by 53% and raised decode throughput up to 1.77 times on a 16-chip TPU v7x slice compared with a similar H200 GPU cluster; the underlying setup and full kernel details sit in a linked write-up not reviewed here.
- mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
- token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
- benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Checking sign-in…
Loading comments…






