A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
Multimodal models map other media into the same internal representation as language, so you can hand them a screenshot, a whiteboard photo, or a call recording and converse about it. For builders the everyday winners are screenshot-to-bug-report, UI-mockup-to-code, and document understanding that reads the actual PDF layout.
Capability is uneven across modes — frontier models read images far better than they generate them, and video understanding still lags — so "multimodal" on a spec sheet is where evaluation starts, not ends.