Vibeleaderboard
← All Intel
Intel / post

we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the seri

Source
wafer_ai
Date
wafer_ai@wafer_ai

we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread 🧵 part 6: Transformer Inference Arithmetic Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s. Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation: - prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase. - batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using…

Read the full post on X
Key takeaways · AI-distilled
  • Building on Carol Chen's (kipply) arithmetic, Wafer notes prefill processes prompt tokens together for more weight reuse, while small-batch decode can spend more time loading weights than computing, so each phase should be profiled separately.
  • Larger batches raise arithmetic intensity in projection and feed-forward matmuls, but they also grow KV-cache requirements and time between output tokens, so the thread advises tuning batch size against a latency target with representative prompt and output lengths.
  • KV-cache capacity and bandwidth need separate accounting: with full , decode reads every retained key and value each step, so a long can fit in memory while its reads dominate latency. Under grouped-query attention, size by KV-head count.
  • More tensor-parallel shards cut per-GPU compute but make NVLink communication a larger share of latency. The thread recommends checking estimates against measured kernel durations and rejecting throughput gains that violate the latency target.
Terms in this piece · Glossary
  • transformer — The neural network architecture behind modern AI models, built on attention — letting every word directly consider every other word in parallel.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters

Explains how to reason about prefill versus decode costs, batching tradeoffs, and memory budgets when sizing inference on H200, B200, and B300 GPUs.

More from wafer_ai
Recommended reads
Comments

Checking sign-in…

Loading comments…