
Had a 💥blast building this with the SGLang team, the collaboration shows what scenario-specific optimization open-source can do. The Humming + DSpark optimizations are already upstreamed. 📖 Full breakdown: https://t.co/QApu2aei3F https://t.co/1Y08FkJyBi
A 1.6T reaching 271 output /s at batch size 1 on H20 hardware without native FP4 tensor cores narrows the gap to B300 to 1.42x, and the Humming and DSpark optimizations behind it are upstreamed in SGLang.
Checking sign-in…
Loading comments…