Introducing Muse Voice Transcribe, the first real-time audio perception model from Meta Superintelligence Labs. Muse Voice Transcribe delivers real-time streaming ASR, diarization with 20+ speakers, and endpointing. It’s multilingual with seamless code-switching and improves accuracy with language, keyword, and context biasing. The model ranks first on @ArtificialAnlys streaming speech-to-text and on public diarization benchmarks.



Muse Voice Transcribe is an autoregressive multimodal LLM from the Muse Spark family. Audio is processed in 80ms chunks (12.5 Hz), one token each, and at every chunk the model decides whether to keep listening or emit text. RL with combined word error rate and delay rewards gives it adaptive delay: it waits longer on hard words and commits sooner on easy ones, trading accuracy against latency word-by-word.
With adaptive delay, Muse Voice Transcribe achieves the pareto frontier on speed-accuracy trade-off measured by time to final transcription.

Muse Voice Transcribe is available today via Meta Model API, Meta AI for Mac, and Muse Code. Learn more: https://t.co/lLNaPF8DlC
transcription with per-word adaptive latency, 20+ speaker diarization, and code switching is available through the Meta Model API, giving voice builders a tunable speed versus accuracy trade-off.
Checking sign-in…
Loading comments…