Haotian Zhang, Fengyuan Yu, Dezhi Luo, Haoran Sun, Zehong Zhao, Qingying Gao, Yihan Li, Siyuan An, Huayi Qin, Yilan Zhang, Zhengze Jiang, Pinyuan Feng, Renrui Zhang, Ziyu Guo, Letian Wang, Mengyue Yang, Kangfu Mei, Maijunxian Wang, Ran Ji, Vikash Kumar, Freda Shi, Chandra Sripada, Vincent C. Muller, Philip Torr, Alan Yuille, Nikolaus Kriegeskorte, Felix Juefei-Xu, Lvmin Zhang, Jieneng Chen, Yilun Du, Hokin Deng
Author
Haotian Zhang, Fengyuan Yu, Dezhi Luo, Haoran Sun, Zehong Zhao, Qingying Gao, Yihan Li, Siyuan An, Huayi Qin, Yilan Zhang, Zhengze Jiang, Pinyuan Feng, Renrui Zhang, Ziyu Guo, Letian Wang, Mengyue Yang, Kangfu Mei, Maijunxian Wang, Ran Ji, Vikash Kumar, Freda Shi, Chandra Sripada, Vincent C. Muller, Philip Torr, Alan Yuille, Nikolaus Kriegeskorte, Felix Juefei-Xu, Lvmin Zhang, Jieneng Chen, Yilun Du, Hokin Deng
Date
Key takeaways · AI-distilled
WROP is 150 hand-designed tasks inspired by cognitive science, grouped into six categories. Blender generators randomize speed, lighting, camera angle and other nuisance parameters while keeping each task's structure, yielding 10,000+ samples per task.
The release includes a 1.5M-sample training corpus and a 300-question exam, on which the authors evaluated 14 video models: 3 reference-to-video, 7 edit and 4 continuation models.
Their 16B world model, PWM-WROP, ranked first among continuation models and third overall in a blind pairwise Elo study, behind only a statistical tie between two reference-to-video models.
Data, exam, model answers, scores and weights are released, along with PWM, a native-PyTorch training stack for AWS Trainium2.
Why it matters
Gives a controlled way to check whether a video/world model actually understands that objects persist and stay solid, a core physical-reasoning gap relevant to anyone building on world models for planning or simulation.