
NVIDIA Research 🚀 has produced some great research, like LatentMoE (used in Kimi K3) and GatedDeltaNets (used in Qwen). But for e2e frontier training, NVIDIA's bureaucratic culture has produced embarrassing models like Nemotron3 Ultra. Despite NVIDIA Research having amazing talent, Nemotron3 Ultra, with 550B total params (55B active), is getting mogged by all the Chinese models, including even Qwen3.8 27B parameters, which has ~20x fewer parameters.

NVIDIA's research shows up inside Kimi K3 and Qwen while its own 550B parameter Nemotron3 Ultra trails a 27B Qwen model, a reminder that research strength does not transfer to end to end training and that parameter count is a poor proxy for capability.
Checking sign-in…
Loading comments…