If you're serving speech-to-text at scale, this breaks down how treating ASR as an end-to-end systems problem — not just GPU kernel tuning — produces state-of-the-art latency and throughput on Artificial Analysis benchmarks.
Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.
Transcript
Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.