eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
An open dataset and scoring methodology for measuring whether a model treats opposing positions with equal depth, combined with a paired harmful/benign prompt design that keeps refusal rates honest. Both are reusable on other models.