Vibeleaderboard
Index / tool
Visit microsoft.github.io
Category
AI Tools
Pricing
Open Source
Type
TOOL
Builder
@microsoft
Added
Apr 12, 2026

About

Open-source voice AI framework that includes advanced speech recognition (ASR) for 60-minute audio transcription with speaker diarization, text-to-speech (TTS) for 90-minute multi-speaker synthesis, and real-time streaming TTS. Operates at ultra-low 7.5Hz frame rate for efficient long-form audio processing.

Why it made the leaderboard

Microsoft Research's open-source voice family built on 7.5Hz continuous speech tokenizers: ASR that transcribes 60 minutes in a single pass with speaker diarization, TTS that sustains 90-minute multi-speaker synthesis, and real-time streaming TTS. Long-form audio without the chunking hacks.

Tags

voice-aispeech-recognitiontext-to-speechopen-sourcelong-form-audiospeaker-diarizationstreaming-ttsmicrosoft

Tech Stack

Python

Featured in Intel

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.