
Agents run dozens of inference calls per task. Latency compounds. Cost compounds. We tested Mercury 2 on @pinchbench — the benchmark built on @openclaw, the fastest-growing open-source project in GitHub history. 78% success rate. Fastest in class. Under $1/M tokens. This is what production agents need 👇

If you run loops, a model at 78% success on an agent for under $1 per million changes the cost of the calls between the interesting ones. The numbers are vendor-run, so check them against your own task mix.
Checking sign-in…
Loading comments…