
Per-dollar and per-watt deltas of this size reset the build-versus-rent calculation.
“Vera Rubin NVL72 running DeepSeek R1 delivers 5.4x performance per MW and 5x performance per dollar over GB200 NVL72 today, and the gap is even wider against GB200 NVL72 during its early bringup in 2025.”
“Blackwell was not able to reuse Hopper WGMMA kernels, Rubin is able to reuse Blackwell's kernels, which makes the software bring up process much smoother.”
“Nvidia shipped 2:4 on weights back in Ampere and nobody used it, because it meant pruning the model and retraining. Rubin applies it to activations at runtime, so no retraining is needed.”
“Rubin is a little over 2x cheaper at low speeds, peaks near 8x at 150 tok/s/user, and then moves back to 5x by 200 tok/s/user.”
Checking sign-in…
Loading comments…