
🚀 AuK is officially here. Nano banana🍌 for audio An open-source foundation model for unified speech generation and editing. Natural-language instructions + reference audio. One interface. Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-accent. Timbre/style/emotion edit. Speed/Pitch control. Enhancement, denoising, multi-speaker and music separation. Also releasing AuK-Flash: 4-step inference. ~4.5× faster under matched conditions. Code, weights, and demo are live. Try it and share your feedback. 🤗 Paper & upvote: https://t.co/zEveUsuJRF ⭐ GitHub & star:
AuK is a single open-weights model covering TTS, voice/style editing, and audio separation, with a 4.5x-faster variant and public code/weights for engineers to self-host.
Checking sign-in…
Loading comments…