
Introducing the next evolution in LLM text generation: the first fully-fledged serving system built around a decode megakernel. Delivering up to 1.58x faster performance than vLLM. Built for North Mini Code, completely open-source.
Gives teams an open-source, faster serving path (up to 1.58x vLLM) built around kernel fusion, with published numbers on H100 hardware.
Checking sign-in…
Loading comments…