
A named, working cascade for fully local speech-to-speech — llama.cpp with Gemma 4, Silero VAD, Parakeet-TDT STT, Qwen3-TTS — behind a Realtime-API-compatible /v1/realtime endpoint, so components can be swapped as new models land.
Checking sign-in…
Loading comments…