Today, MiMo can see We release MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning. MiMo-VL-7B-RL outperforms Qwen2.5-VL-7B on 35 out of 40 evaluated tasks, and scores 59.4 on OlympiadBench, surpassing even Qwen2.5-VL-72B and GPT-4o . For GUI grounding applications, it sets a new standard with 56.1 on OSWorld-G, even outperforming specialized models such as UI-TARS. To facilitate the evaluation of multimodal tasks, we release a comprehensive evaluation suite covering over 50 tasks to promote reproducibility and advance the field, which will be open-sourced soon. Our paper, model and code can be accessed via the following link. Github: https://t.co/aatkHFZ7Oz Huggingface:

What can MiMO-VL-RL do? 🤔 In Case #1, our model showcases strong plot understanding capabilities, successfully converting an intricate plot into a well-structured markdown table. It also demonstrates superior reasoning capabilities in STEM tasks. In Case #2, MiMo-VL-7B successfully navigates a webpage to complete the task of purchasing a Xiaomi SU7 with customized paint and interior options.


The comprehensive visual perception capabilities of MiMo-VL-7B are attributed to high-quality pre-training data and an innovative Mixed On-policy Reinforcement Learning (MORL) algorithm: Multi-Stage Pre-training: High-quality pre-training multimodal data has been collected, cleaned, and synthesized, covering data types such as image-text pairs, video-text pairs, and GUI operation sequences, totaling 2.4T tokens. By adjusting the proportions of different types of data in stages, it enhances the ability for long-range multimodal reasoning. Mixed Online Reinforcement Learning: It combines various feedback signals including mixed text reasoning, multimodal perception + reasoning, and RLHF, and employs online reinforcement learning algorithms to stabilize and accelerate training, significantly improving the model's reasoning, perception performance, and user experience.

On general benchmarks, the MiMo-VL-7B models, particularly MiMo-VL-7B-SFT and MiMo-VL-7B-RL, demonstrate consistently leading performance across a diverse range of vision-language benchmarks, surpassing other open-source models of comparable or larger scale.

A 7B general VLM scoring 56.1 on OSWorld-G suggests computer-use agents may not need a GUI-specialist model, which removes a component from the stack.
postHeads up, agent users! If you're using Xiaomi MiMo with thinking mode: When…
postMiMo-V2.5 and V2.5-Pro go open weights under MIT with day-zero SGLang and vLLM
postIntroducing MiMo-V2.5 Voice — our full-stack voice lineup for the Agent era. 🚀…Checking sign-in…
Loading comments…