1/ The fastest reasoning LLM is now live exclusively on OpenRouter. Mercury 2.5 Preview from @_inception_ai reaches 1,107 tokens/sec through parallel token generation, with tunable reasoning, parallel tool calls, and schema-aligned JSON. Built for latency-sensitive workloads.
2/ Use it now: https://t.co/XhEeymoGbi
If throughput is the bottleneck, a quoted at 1,107 per second with parallel tool calls and schema-constrained JSON sits at a very different point on the latency curve than autoregressive models.
Checking sign-in…
Loading comments…