Vibeleaderboard
← All Intel
Intel / post

StepFun's StepAudio 3 ASR Tops Speech-to-Text Accuracy Leaderboard

Source
ArtificialAnlys
Date
ArtificialAnlys@ArtificialAnlys

StepFun has released StepAudio 3 ASR, ranking #1 on the AA-WER Index for non-streaming Speech to Text with 1.7% WER, a notable improvement on StepAudio 2.5 ASR (4.7%) StepAudio 3 ASR is StepFun's new Speech to Text model, available through the StepFun API for non-streaming transcription, and the first StepFun model to reach the top of our non-streaming Speech to Text leaderboard. Key takeaways ➤ Accuracy: StepAudio 3 ASR scores 1.7% on the AA-WER Index, effectively tied with Alibaba's Fun-Realtime-ASR-preview at 1.7%. It leads on AA-AgentTalk, our held-out voice-agent dataset, at 1.4%, but trails on long-form Earnings22 calls at 2.8%, where Fun-Realtime-ASR-preview scores 1.8% ➤ Speed: The model transcribes at a speed factor of 88x real time, behind MAI-Transcribe-2 at 374x, Smallest AI Pulse Pro at 285x and Grok Voice Transcribe 2.0 at 154x, roughly level with ElevenLabs Scribe v2 at 84x and Gemini 3.5 Transcribe at 91x, and ahead of Fun-Realtime-ASR-preview at 20x ➤ Price: StepAudio 3 ASR costs $0.40 per hour, or $6.67 per 1,000 minutes, the most expensive of the five most accurate models. MAI-Transcribe-2 and Grok Voice Transcribe 2.0 cost $1.67 per 1,000 minutes and…

Read the full post on X

Context

Speech-to-text accuracy is usually measured by word error rate (WER), the share of words a transcription model gets wrong compared with a human transcript. Artificial Analysis reports that StepFun's new StepAudio 3 ASR model now tops its non-streaming Speech-to-Text leaderboard with a 1.7% WER, a sharp improvement over StepFun's own prior model, StepAudio 2.5 ASR, which scored 4.7%, and the first time a StepFun model has led that leaderboard.

That top score is effectively tied with Alibaba's Fun-Realtime-ASR-preview, also at 1.7%. Looking at specific conditions, StepAudio 3 ASR leads on Artificial Analysis's held-out voice- dataset at 1.4% WER, but trails on long-form earnings-call audio at 2.8%, where Fun-Realtime-ASR-preview scores 1.8%, meaning its accuracy advantage is not uniform across audio types.

The tradeoff is speed and price: StepAudio 3 ASR transcribes at about 88x real time, slower than several rivals like MAI-Transcribe-2 at 374x, and costs $0.40 per hour of audio, the most expensive of the five most accurate models Artificial Analysis tracks. The reported takeaway is that it buys top-tier accuracy on conversational audio specifically, at a higher per-minute cost than faster competitors.

Terms in this piece · Glossary
  • streaming — Sending a model's response token by token as it is generated, so the reader sees text immediately instead of waiting for the whole answer.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…