We’re excited to launch Mercury, the first commercial-scale diffusion LLM tailored for chat applications! Ultra-fast and efficient, Mercury brings real-time responsiveness to conversations, just like Mercury Coder did for code.
According to 3rd-party benchmarking from @ArtificialAnlys, Mercury matches the performance of speed-optimized frontier models like GPT-4.1 Nano and Claude 3.5 Haiku while running over 7x faster.

Mercury’s low latency enables it to power responsive voice applications, ranging from translation services to call center agents. On real-world voice prompts and standard Nvidia hardware, Mercury provides lower latency than Llama 3.3 70B running on Cerebras.

Mercury is also the founding LLM partner for @Microsoft NLWeb project. Compared with other speed-focused models like GPT-4.1 Mini and Claude 3.5 Haiku, Mercury runs far faster, ensuring a fluid user experience. https://t.co/g647IORy8y

Mercury brings diffusion-based parallel generation to chat at over seven times the speed of comparable small frontier models, with third-party benchmarks and immediate access via API, OpenRouter, and Poe, which makes it viable for voice and other latency-bound applications.
Checking sign-in…
Loading comments…