We're releasing MolmoMotion, a 3D motion forecasting model. Given one or a few video frames, 3D points on an object, & an instruction like "Put the white bowl on the table," MolmoMotion predicts where those points will go over the next few seconds in a shared 3D world frame. 🧵
MolmoMotion can predict different motions across scenes, like a bowl sliding and rotating on a table, a flamingo dipping its beak as it walks, & a lint roller working back and forth on cloth. Each predicted path follows the instruction and stays close to the ground-truth motion.
MolmoMotion represents motion as 3D points attached to an object, tracked in a shared world frame. This approach doesn't need templates for objects, stays stable as camera perspectives change, & is compact enough to feed straight into downstream applications.

Motion forecasters like MolmoMotion have many potential applications. Fine-tuned, MolmoMotion can predict object paths to help grasping robots plan where to move objects, or help image generators capture motion more accurately—particularly motions hard to describe in a prompt.
MolmoMotion forecasts where points on an object will move over the next few seconds from one or a few frames plus an instruction, without per-object templates, and ships with a 1.16M-video dataset and a for anyone building grasp planning or motion-aware generation.
Checking sign-in…
Loading comments…