eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
Frames foundation model training as an automation problem rather than a staffing one, and enumerates the six capabilities a lab needs to iterate faster than it can hire: automated evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition →, RL from execution, architecture ablations, and data mixing among them.