🚀 1,000+ TOKENS/S ON A 1T MODEL! 🚀 We are thrilled to release Xiaomi MiMo-V2.5-Pro-UltraSpeed in collaboration with @TileRT_AI , breaking the 1,000 tokens/s output speed on a 1 Trillion parameter model for the FIRST TIME! Not wafer-scale integration like Cerebras. Not pure on-chip SRAM chips like Groq. We achieve 1,000 tps on a 1T MoE model using just a SINGLE, STANDARD 8-GPGPU NODE. Read the full technical deep dive:https://t.co/MX0kjHKdKi Want to experience the future of real-time AI? 👉 Apply for UltraSpeed now: https://t.co/aeWAxyhwVk ⏳ Limited-Time Access: Application-based · Jun 8 – Jun 23 (PDT) 💬 Chat Experience: Completely FREE for a limited time — try the blazing-fast web chat now. ⚡ UltraSpeed API: Just 3x the price for a ~10x boost in output experience. 🤝 Enterprise & Large-Scale Needs: business-mimo@xiaomi.com

🔓 And the best part — we're open-sourcing it. 1,000+ tps on a 1T model wasn't a single breakthrough — it's deep model × system co-design between the MiMo and TileRT teams, all on general-purpose GPUs (no Cerebras-style wafer-scale, no Groq-style SRAM ASICs). On the model side: FP4 quantization (smaller footprint, less memory traffic) + DFlash, our block-masked parallel speculative decoding that accepts far more tokens per verification. On the system side, TileRT tailors its compiler & kernels to exactly these techniques. The result: a 1T model breaking 1,000 tps on a single, standard 8-GPU node. 🤗 Open weights (FP4 + DFlash checkpoint): https://t.co/jYQsgeruMg
Shows frontier-scale decode speed is reachable on commodity 8-GPU nodes through FP4 plus , which changes the latency you can assume for large-model loops without specialty silicon.
postHeads up, agent users! If you're using Xiaomi MiMo with thinking mode: When…
postMiMo-V2.5 and V2.5-Pro go open weights under MIT with day-zero SGLang and vLLM
postIntroducing MiMo-V2.5 Voice — our full-stack voice lineup for the Agent era. 🚀…Checking sign-in…
Loading comments…