
The fastest reasoning LLM is now in production on Baseten. Mercury 2 is a diffusion LLM, so it generates tokens in parallel and hits 1,000+ tokens/sec on @NVIDIAAI GPUs, speeds that used to require specialized hardware. @augmentcode is already using Mercury 2, cutting cost 90% and latency 82%. Proud to partner with the @baseten team to bring dLLMs to production.
Another production serving path for a diffusion , with a named customer reporting 90% lower cost and 82% lower latency after switching workloads to Mercury 2.
Checking sign-in…
Loading comments…