
Mercury 2.5 from @_inception_ai is now generally available on OpenRouter. It's the fastest model on our Fastest Models leaderboard that doesn't require specialized hardware. A diffusion LLM: it decodes tokens in parallel rather than one at a time. The gain is largest on code-heavy prompts. Use it here: https://t.co/xTId0uFKC8
Diffusion-based decoding offers a genuinely different latency profile than autoregressive models, worth testing for code-generation workloads where -by-token decoding is the bottleneck.
Checking sign-in…
Loading comments…