
One model generates up to 20 seconds of video with native synchronized audio, conditioned on text, images, or reference clips. The same backbone also predicts robot actions, and the developers argue that joint training beats per-modality specialists.
Checking sign-in…
Loading comments…