Physicalrealismbench Attributable Physical Realism Evaluation For Video World Models
Source
Reka AI editorial sitemap
Author
Reka AI editorial sitemap
Date
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
If you use a VLM as a judge or reward function for world models, this work shows it may miss objects vanishing outright, and that binary violated/not-violated scoring cannot distinguish perception from a lucky guess.