MiniMax reports that its SGLang-Diffusion sampler, VDN-H3, now generates 14.4 seconds of 768p video in 9.0 seconds end to end on eight B200 GPUs, more than twice real time. The company says the speedup comes with no measured quality regression against the 50-step dense baseline it was tested against, meaning the gain is from a better sampling schedule rather than a smaller or cheaper model. Faster inference changes the economics of video generation directly, since compute time per clip is most of the cost at scale. It also demonstrates that open inference stacks and open model weights are compounding gains together rather than independently, since the sampler improvement applies on top of whatever model efficiency MiniMax already had. Teams running video generation in production get a concrete lever to cut GPU hours per video without waiting for a new model release.
H3 keeps getting faster. ⚡️ @sgl_project + VDN-H3 now push MiniMax H3 beyond 2× real-time denoising on 8× B200 - generating 14.4s of 768p video in 9.0s end-to-end after warmup, with no measured quality regression. Open models compound through open ecosystems. 🚀
Checking sign-in…
Loading comments…