multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
Why it matters
Explains why computer use agents need both pixels and structured state: DOM and accessibility trees make runs cheaper and debuggable, while vision catches UI states that exist only visually.