Cache-aware prefill–decode disaggregation (CPD) for up to 40% faster long-context LLM serving
Source
www.together.ai
Date
Why it matters
If you're serving long-context LLMs and fighting slow time-to-first-token, CPD shows how separating cache-warm and cache-cold workloads across prefill and decode stages can lift throughput ~40% — a concrete architectural lever beyond generic prompt caching.
Serving long prompts doesn't have to mean slow responses.
Learn how Together AI's CPD architecture separates warm and cold inference workloads to deliver 40% higher throughput and dramatically lower time-to-first-token for long-context LLM serving.
Transcript
Serving long prompts doesn't have to mean slow responses. Learn how Together AI's CPD architecture separates warm and cold inference workloads to deliver 40% higher throughput and dramatically lower time-to-first-token for long-context LLM serving.