Mercury 2 is live. The world's first reasoning diffusion LLM – 5x faster than leading speed-optimized autoregressive models. Built for production: multi-step agents without delays, voice AI with tight latency budgets, instant coding feedback. Diffusion-based generation enables parallel refinement, not sequential tokens. Faster. More controllable. Dramatically lower inference cost. Available today on the Inception API. @dinabass has the story in @business.

https://t.co/2WTHJUfs4a
A production diffusion with reasoning gives a second architecture to reach for when latency is the constraint: multi-step agents, voice, and inline coding feedback where sequential decoding is too slow.
Checking sign-in…
Loading comments…