
🎤 Introducing Step-Audio-R1.1: The New Frontier of Audio Reasoning! 🏆 We just hit No.1 on the Artificial Analysis Speech Reasoning leaderboard! Our results: ✅96.4% accuracy on BigBench Audio, setting a new record and surpassing Grok, Gemini, OpenAI, and Google models (Fig. 1). ✅1.51s TTFA (Time-to-First-Audio), fast enough to feel like a real conversation 🚀Step-Audio-R1.1 demonstrates that deep reasoning and real-time performance are no longer a trade-off. As shown in Fig. 2, it maintains the highest reasoning depth in class while staying under the critical ~1.5s latency threshold. What’s under the hood: ✔️Test-time compute scaling: Step-Audio-R1, the first audio LLM to unlock dynamic compute allocation during inference. ✔️End-to-end audio reasoning: coherent, on-the-fly reasoning with zero added latency. ✔️Scalable CoT optimized specifically for audio understanding tasks. R1.1 is the upgraded, smarter, faster evolution. Weights are OPEN. Go build something crazy. 🤗 HuggingFace: https://t.co/VZAKvk7v1h 🎤 Try it here: https://t.co/g3v0fGjX5y 🔮 ModelScope:


Voice agents usually trade reasoning depth against response time. This model reports leading BigBench Audio accuracy while staying near 1.5s to first audio, and the weights are open so you can test that claim.
Checking sign-in…
Loading comments…