eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
Parallel sampling plus a trained resolver buys 3-4 points on research benchmarks at near-flat latency, and the low/high API modes let you price the accuracy tradeoff directly at $35 vs $60 per 1k requests.