Etalon: measuring LLM streaming latency beyond TPOT averages
- Source
- wafer_ai
- Date
this is a must-read if you want to be a top 1% ai performance engineers follow and save to keep up with the series. links in thread 🧵 part 8: Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems Amey Agrawal and coauthors explain how serving decisions shape the timing of streamed LLM responses. for AI performance engineers evaluating batching, chunked prefill, or speculative decoding, the paper provides a framework for assessing token delivery against application-specific timing targets. the authors connect scheduling behavior, token arrival traces, and service targets: - prefill and decode interference: a long prefill can delay ongoing decode work. chunked prefill divides that prompt processing across batches, so the amount of prefill work per batch can affect both first-token latency and interruptions to existing streams. - first-token latency: time to first token (TTFT) includes scheduling delay and prompt processing. the authors propose profiling isolated requests across prompt lengths and adding a scheduling allowance to set first-token deadlines. - averages and latency distributions: time per output token (TPOT) can hide generation pauses,…

- Chunked prefill spreads a long prompt's processing across batches, so the amount of prefill work packed into each batch shifts both time to first and the interruptions felt by streams that are already decoding.
- Because TTFT includes scheduling delay as well as prompt processing, the authors propose setting first-token deadlines by profiling isolated requests across prompt lengths and then adding a scheduling allowance.
- Etalon's fluidity-index scores the share of per-token deadlines a request meets: early tokens bank slack, and after a miss it counts the missed slots and resets later deadlines from that token's actual arrival.
- can deliver several tokens in one burst; the thread notes a client can buffer them for steady display, so work finished ahead of the reading rate covers some later gaps.
- For capacity planning, Etalon searches for the highest request rate one replica can sustain under chosen targets on the tested workload; the thread advises fixing model, hardware, workload and targets when comparing serving configs.
- LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
- inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
- speculative decoding — A speed trick where a small model drafts several tokens ahead and the big model verifies them in one pass, often doubling generation speed.
- token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Explains why average TPOT and TBT percentiles hide generation stalls and how per-request token arrival traces expose them. Useful when tuning batching, chunked prefill or speculative decoding.
postInference scaling checklist drawn from the PaLM TPU v4 paper
postwe launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the seri
postWafer says it beat Cerebras on latency for YC's AI Office Hours
postWafer maps a classic GPU textbook onto real AI performance engineering work
Checking sign-in…
Loading comments…