Physicalrealismbench Attributable Physical Realism Evaluation For Video World Models
Source
Reka AI editorial sitemap
Author
Reka AI editorial sitemap
Date
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
VLMs are increasingly used as judges and reward functions for world models, and this benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → shows they miss basic physics violations while scoring well on binary verdicts, producing a false sense of progress in any evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → built on them.