Introducing MiMo-V2.5 Voice — our full-stack voice lineup for the Agent era. 🚀 Voice AI is moving beyond just hearing and reading — towards precise understanding and flexible expression. That’s why today we’re launching the MiMo-V2.5-TTS Series and open-sourcing MiMo-V2.5-ASR. The TTS series includes three models: 🔹 MiMo-V2.5-TTS — premium built-in voices with fine-grained style control 🔹 MiMo-V2.5-TTS-VoiceDesign — generate brand-new voices from natural-language descriptions 🔹 MiMo-V2.5-TTS-VoiceClone — clone target voices with high fidelity from just a few samples On the recognition side, MiMo-V2.5-ASR is now open source, with strong performance across bilingual speech, Chinese dialects, code-switching, noisy audio, and multi-speaker scenarios. How to get started: For MiMo-V2.5-TTS Series: • It is available for a limited time for free on the Xiaomi MiMo API platform: https://t.co/6hWRgaz4yP • You can also try the TTS models instantly in Xiaomi MiMo Studio: https://t.co/VGlPAi2va9 • See more cases at: https://t.co/e2nYhnGpvV • We’ve also open-sourced MiMo TTS Skills for fast integration into Agent applications: https://t.co/uPOYhASCpC For MiMo-V2.5-ASR: • Demo: https://t.co/ju5aXdNDBD • Code: https://t.co/jpC5T2zOQi • Weights: https://t.co/z8CD9F2Lgu • Hugging Face Space: https://t.co/xjlL0E5TJn Jump in and start building with MiMo-V2.5 Voice today!

What makes the MiMo-V2.5-TTS Series different? It’s not just about generating speech. It’s about making voice controllable, expressive, and usable in real creative workflows. Across the series, the models share three core strengths: 🔹 Strong instruction following — from a single prompt to a full director’s note, they can reliably follow guidance on emotion, tone, pace, delivery, and style 🔹 Flexible audio tag control — inline audio tags let you precisely control emotion, state, and style at specific points in the text, from simple cues to complex multi-tag combinations within the same passage 🔹 Rich text understanding — even with plain text only, the models can naturally pick up rhythm, pauses, emotional shifts, and character cues On top of that shared foundation, each model is built for a different use case: MiMo-V2.5-TTS High-quality built-in voices, professionally tuned for natural pronunciation and expressive delivery — ready to use out of the box. MiMo-V2.5-TTS-VoiceDesign Generate entirely new voices from natural-language descriptions — no reference audio needed. From age and accent to texture, temperament, and speaking style, the model can turn rich descriptions into distinctive voices, including character voices that don’t exist in any preset library. MiMo-V2.5-TTS-VoiceClone With just a few seconds of reference audio, the model can clone a target voice with high fidelity, with no extra training or fine-tuning. It preserves not only the speaker’s identity, but also personal details like breath, rhythm, and pauses — while still supporting the full control stack of the series. Whether you need a ready-made voice, a brand-new one, or your own cloned one, MiMo-V2.5-TTS gives you the same thing underneath: voice that can actually follow direction. https://t.co/y5KOMI2GE5
MiMo-V2.5-ASR — hear every word, even when speech gets messy. Good ASR isn’t just about transcribing clean audio. It’s about understanding real speech: accents, code-switching, noisy environments, overlapping speakers, and dense knowledge content. As the listening foundation of the MiMo-V2.5 voice lineup, MiMo-V2.5-ASR delivers strong real-world recognition across Chinese and English, including: Chinese dialects — Wu, Cantonese, Minnan, Sichuanese, and more Complex English audio — strong results in challenging scenarios like AMI Code-switching — fluid Chinese-English transcription with no language tags required Lyrics recognition — accurate song transcription in Chinese and English Noise robustness — reliable recognition in high-noise and far-field conditions Multi-speaker scenarios — accurate transcription for meetings and overlapping dialogue Knowledge-heavy speech — better handling of poems, technical terms, names, and locations Native punctuation — punctuation generated directly, with no post-processing needed The results below show the pattern clearly: MiMo-V2.5-ASR is state-of-the-art or highly competitive across Chinese, English, dialect, code-switch, and lyrics recognition — with strong consistency across languages and real-world settings.

Open ASR weights plus TTS with inline per-span style tags gives a self-hostable listening layer and directable speech output, rather than one fixed voice behind an API.
Checking sign-in…
Loading comments…