Vibeleaderboard
Index / tool

MiMo-V2.6-Flash-MOPD

huggingface.co
Visit huggingface.co
Category
Developer Tools
Rank
No. 2835Tools index
Pricing
Open Source
Type
TOOL
Date

About

MiMo-V2.6-Flash-MOPD is a multimodal large language model from Xiaomi's MiMo team that handles text generation, vision-language understanding, audio, video understanding, and agentic tool calling with long-context support, released under the MIT license. Its architecture combines an LLM backbone with a dedicated MiMo vision encoder, audio encoders, and a speculative decoder, and it can be served through vLLM or SGLang for OpenAI-compatible inference.

Why it made the leaderboard

Gives developers a lower-latency, MIT-licensed option for running Xiaomi's multimodal agent model locally or via popular inference servers.

Tags

llmmultimodalvision-languageaudioai-agentsopen-source

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.