eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
alignment — The work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.
Why it matters
A concrete step toward externally verifiable AI safety oversight: embedded evaluators get access to training decisions and staff, not just black-box outputs, changing how frontier model safety claims can be checked.