eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
Closes the robot-learning loop inside one library: policies that predict future frames, reward models that score success, and a rollout CLI that turns deployment failures into training data.