eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
The ASL framework is the stated condition under which a frontier lab withholds deployment or pauses training, so it sets the terms on which future model capability and availability actually arrive.