
If you're serving speech-to-text at scale, this breaks down how treating ASR as an end-to-end systems problem — not just GPU kernel tuning — produces state-of-the-art latency and throughput on Artificial Analysis benchmarks.
“The faster of the two, NVIDIA Parakeet-TDT 0.6B v3, can transcribe roughly 20 hours of speech, about the runtime of the Harry Potter film franchise, in under 10 seconds.”
Together AI
“The same Harry Potter corpus as audiobooks is 5 to 10 GB, roughly three orders of magnitude larger than the text.”
Together AI
“The CPU leaves the decoder’s inner loop, and the result is a 2 to 3x faster decoder.”
Together AI
“Under load for streaming workflows, p50 and p90 latency looked healthy, but p95 would periodically spike by about 200 ms.”
Together AI
“The lesson was to keep profiling beyond the model. GPU time, queue depth, and model execution all looked normal; the latency spike lived in the Python runtime.”
Together AI
Checking sign-in…
Loading comments…