We tested Mercury Decide, a new model from @_inception_ai, on the same benchmark we used for TypeSafe's Jev.
It matched frontier accuracy on claim verification at the lowest cost we have measured so far.
Shows a diffusion-based decision model matching frontier accuracy on claim verification at a quarter of Jev's cost, with calibrationHow well a model's confidence matches reality — a calibrated model saying "90% sure" is right about 90% of the time.Full definition → confidence for routing. Latency grows with batched questions, so check your workload shape.
Terms in this piece · Glossary
calibration — How well a model's confidence matches reality — a calibrated model saying "90% sure" is right about 90% of the time.