
fusion-runtime
github.com/samarthurs18/fusion-runtimeAbout
fusion-runtime is a self-hosted voice agent runtime that streams speech-to-text, an LLM, and text-to-speech into each other within a single process, so a reply starts playing before generation finishes. On an RTX 3090 running a 7B model it reports about 490ms of processing after a turn ends (991ms from the user's last syllable) at roughly 127 tokens/sec, and it supports mid-sentence interruption.
Why it made the leaderboard
Gives builders a self-hosted alternative to hosted voice-agent APIs with real performance data and mid-sentence interruption handling for latency-sensitive local deployments.
Tags
voice-agentself-hostedspeech-to-texttext-to-speechlow-latency
Tech Stack
Python
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.