
Today, we’re introducing Mercury 2.5, the most capable diffusion LLM on the market. It offers a 40% jump in intelligence over Mercury 2, and runs over 1,100 tokens/sec on widely-available @NVIDIAAI GPUs. It’s available today on our API, @OpenRouter, and @Baseten. Contact us to evaluate Mercury 2.5 for production:
A diffusion claiming over 1,100 /sec on standard NVIDIA hardware gives latency-sensitive applications a faster alternative to autoregressive models, now accessible via API, OpenRouter, and Baseten.
Checking sign-in…
Loading comments…