
If you're building or evaluating multimodal applications, this reveals that models can answer the same question far less accurately when it's presented via image instead of text — a hidden failure mode worth testing for before shipping vision-language features.
“we observe a pervasive modality gap, with performance decreasing by 4.2% to 47.4% when moving from text-only to image-interleaved inputs, averaging 19.6% +/- 3.3% across models”
“reasoning models exhibit consistently smaller gaps, achieving an average gap of 10.1% compared to 25.5% for non-reasoning models”
“neither prompting strategies nor scaling training compute alone reliably reduces the modality gap”
articleMultimodality and Large Multimodal Models (LMMs)Chip Huyen
articleCapability-Driven Multimodal Scaling LawZiran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai
articleCompliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System MessagesJuan Yeo, Geewook KimChecking sign-in…
Loading comments…