multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
Why it matters
Gemini Robotics 2 extends multimodalA model that works with more than text — reading images, audio, or video, and sometimes generating them too.Full definition → reasoning into coordinated full-body physical action and transfers across different robot bodies — a marker of where VLA models now sit versus scripted or teleoperated control.