Vibeleaderboard
← All Intel
Intel / video

Agents That Own Their Inference — Du'an Lightfoot & Khaja Omer, Akamai Technologies

Source
youtube.com
Author
AI Engineer
Date
Why it matters

A faster-sounding optimization like speculative decoding can cut throughput from about 58 to 16 tokens per second. The workshop shows how to benchmark FP8, concurrency and sequence limits against your own model and hardware.

Key takeaways · AI-distilled
  • Each participant gets a Kubernetes namespace with a Jupyter notebook and a vLLM endpoint serving a Qwen model, used to measure how an 's prompts, responses and repeated tool calls consume latency and memory.
  • The notebooks separate per-request generation speed from aggregate throughput, compare cold and warm prefixes to show effects, and contrast dense models with mixtures of experts.
  • When slowed the model, Omer examined the drafter's acceptance rate, model compatibility and GPU contention, then returned to the baseline. An earlier FP8 switch was checked against simple output quality cases, not just speed.
  • Final tuning finds the concurrency point where throughput flattens while latency keeps rising, then changes the sequence limit and measures again. The presenters note that a notebook failure changed the model used in that last comparison.
Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • streaming — Sending a model's response token by token as it is generated, so the reader sees text immediately instead of waiting for the whole answer.
  • KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
  • speculative decoding — A speed trick where a small model drafts several tokens ahead and the big model verifies them in one pass, often doubling generation speed.
Read the source www.youtube.com
More from AI Engineer
Recommended reads
Comments

Checking sign-in…

Loading comments…