eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
WorldModelGym reframes world-model evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → around whether a model's predictions support correct decisions, not whether rollouts look realistic. A single interface handles latent, predictive, and pixel models under frozen policies and is cheap enough to run on heavy models.