
Mercury 2 is now available on Azure Foundry. The world's fastest reasoning language model, built to make production AI feel instant. Mercury 2 doesn't decode sequentially. It generates responses through parallel refinement, producing multiple tokens simultaneously and converging over a small number of steps. Over 1,000 tokens per second on widely-deployed NVIDIA GPUs, at less than half the cost, with comparable quality to Claude Haiku 4.5 and GPT-5 Mini. Full announcement:
Azure customers can call a diffusion without leaving the platform, at a claimed 1,000+ per second and under half the cost of comparable small autoregressive models.
Checking sign-in…
Loading comments…