
An open, deterministic training kernel with published speedups lowers the cost floor for teams training their own models rather than only consuming APIs.
“As we have scaled the training and inference of Composer , our agentic coding model, the mixture-of-experts layer has consistently remained the major bottleneck. Depending on the workload and training configuration, it can consume more than half of end-to-end training time.”
cursor_ai
“MoK addresses that bottleneck by fusing all MoE communication and computation into a single, fully deterministic kernel. It now powers Composer training across tens of thousands of GPUs.”
cursor_ai
“Mixture-of-Kittens achieves up to 2.37x higher MXFP8 forward throughput than the fastest public baseline on GB300 NVL72s”
cursor_ai
“In our production training stack across several NVL72 racks, MoK increased end-to-end tokens per second by 1.41x.”
cursor_ai
“Empirically, pull-based communication delivers up to 29% higher NVLink bandwidth utilization than push-based communication when dispatching tokens under expert imbalance, making it the better choice.”
cursor_ai
Checking sign-in…
Loading comments…