
This provides evidence that a video generation backbone absorbs robot action prediction without permanent capacity cost. Quality dipped 10% and recovered in 3,500 steps, supporting the claim that video, audio, and action share one world representation.
Checking sign-in…
Loading comments…