A repeatable set of representative cases used to measure system behavior.
An evaluation dataset should include ordinary tasks, edge cases, prior failures, and adversarial examples from the real distribution. Its labels and rubrics need versioning because the definition of acceptable behavior changes.