
Auditing official V-JEPA checkpoints shows their action ranking by latent distance is unreliable: the top-scored candidate is usually wrong in manipulation tasks, a defect that survives backbone swaps.
articleAutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model ResearchMarjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
articleThe Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory AgentsJundong Hu, Shekar Ramachandran
articleBelief-Calibrated Optimization: An Explicit World Model for Agentic OptimizationYuhan Chen, Zhihua Tian, Mahavir Dabas, Charith Peris, Rahul Gupta, Ming Jin, Feiyang Kang, Siyuan Zhang, Nan Wang, Ruoxi JiaChecking sign-in…
Loading comments…