Agents That Own Their Inference — Du'an Lightfoot & Khaja Omer, Akamai Technologies
Source
youtube.com
Author
AI Engineer
Date
Why it matters
A faster-sounding optimization like speculative decoding can cut throughput from about 58 to 16 tokens per second. The workshop shows how to benchmark FP8, concurrency and sequence limits against your own model and hardware.
Key takeaways · AI-distilled
Each participant gets a Kubernetes namespace with a Jupyter notebook and a vLLM endpoint serving a Qwen model, used to measure how an AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition →'s prompts, responses and repeated tool calls consume latency and memory.
The notebooks separate per-request generation speed from aggregate throughput, compare cold and warm prefixes to show KV cacheThe memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.Full definition → effects, and contrast dense models with mixtures of experts.
When speculative decodingA speed trick where a small model drafts several tokens ahead and the big model verifies them in one pass, often doubling generation speed.Full definition → slowed the model, Omer examined the drafter's acceptance rate, model compatibility and GPU contention, then returned to the baseline. An earlier FP8 switch was checked against simple output quality cases, not just speed.
Final tuning finds the concurrency point where throughput flattens while latency keeps rising, then changes the sequence limit and measures again. The presenters note that a notebook failure changed the model used in that last comparison.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
streaming — Sending a model's response token by token as it is generated, so the reader sees text immediately instead of waiting for the whole answer.
KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
speculative decoding — A speed trick where a small model drafts several tokens ahead and the big model verifies them in one pass, often doubling generation speed.