eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
It gives builders integrating Runway's API real data on how much quality a cost-optimized router sacrifices versus always calling the top model, informing production routing decisions.