StepAudio 2.5 TTS is live now! Control emotion, pacing, pauses, and delivery with plain natural language. No tags, no preset combos. Just describe what you want the voice to do. Zero-shot voice cloning with full timbre + emotion control. Available via Pay-as-you-go API or Step Plan.
Docs: - Standard API: https://t.co/4f2MmO1VOF - Step Plan: https://t.co/b4sev5xWqE
Direct delivery, pauses and emotion with a plain language instruction rather than markup tags, plus zero shot voice cloning that keeps timbre and emotion separately controllable.
article👏🏻Congratulations!Step3-VL-10B was selected for HuggingFace Daily Papers…
postStep-Audio-R1.1 opens weights for a speech model that reasons in real timeChecking sign-in…
Loading comments…