eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
It shows how to make a prompt change fail CI like a broken unit test, including setting the threshold from measured run-to-run noise and avoiding the skipped-workflow trap in required checks.