MiMo-V2.6-Flash-MOPD
huggingface.co- Category
- Developer Tools
- Rank
- No. 2835Tools index
- Pricing
- Open Source
- Type
- TOOL
- Date
About
MiMo-V2.6-Flash-MOPD is a multimodal large language model from Xiaomi's MiMo team that handles text generation, vision-language understanding, audio, video understanding, and agentic tool calling with long-context support, released under the MIT license. Its architecture combines an LLM backbone with a dedicated MiMo vision encoder, audio encoders, and a speculative decoder, and it can be served through vLLM or SGLang for OpenAI-compatible inference.
Why it made the leaderboard
Gives developers a lower-latency, MIT-licensed option for running Xiaomi's multimodal agent model locally or via popular inference servers.
Tags
llmmultimodalvision-languageaudioai-agentsopen-source
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.