multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
Why it matters
Training on temporally dense captions gives fine-grained control over when things happen in a shot, enabling keyframed transitions and expressive human performance the prior generation could not sustain.