
StepAudio 2.5 ASR is here! 4x faster decoding. 30-minute continuous audio. $0.022/hour. We built this to push ASR further — faster inference, longer inputs, bilingual accuracy that holds up, and a price point that actually scales. ⚡ 4x decoding speed vs. previous generation ⚡ 30-min continuous audio processing ⚡ Top-tier bilingual (EN/ZH) performance ⚡ 80% lower inference cost — API at $0.022/hr For more details, check out: - Documentation: https://t.co/oiya8DpmfM - Demos: https://t.co/9ob48LOGqs - Model card: https://t.co/xo1YDteuQG
Transcription at $0.022 per hour with 30 minute continuous audio and 4x faster decoding changes the cost math for bulk English and Chinese speech pipelines.
article👏🏻Congratulations!Step3-VL-10B was selected for HuggingFace Daily Papers…
postStep-Audio-R1.1 opens weights for a speech model that reasons in real timeChecking sign-in…
Loading comments…