Reading the CUDA Programming Guide for LLM inference speed
- Source
- x.com
- Date
we launched the most comprehensive ai performance engineering repo in the world and we're posting every single resource follow and save to keep up with the series. links in thread 🧵 part 10: CUDA Programming Guide making an LLM faster means understanding where execution is waiting. an attention kernel can wait for its next tile of data, while the serving runtime can leave the GPU waiting for the next launch. NVIDIA's CUDA Programming Guide explains the mechanisms behind those waits. for a modern AI performance engineer working in PyTorch, Triton, or CUDA, it helps with: - tuning attention, matmul, and normalization kernels. the SIMT and CUDA Tile chapters connect layouts and data reuse to GPU execution. the resource guidance explains why larger tiles or more fusion can introduce register spills. - overlapping data movement with computation. asynchronous copies, TMA, and producer/consumer pipelines explain how supported GPUs can load the next tile while processing the current one, with the synchronization required to reuse buffers safely. - managing growing KV caches. the virtual memory chapter explicitly connects its allocation model to LLM serving: reserve an address…

Connects specific CUDA features (TMA pipelines, virtual memory for growing KV caches, CUDA Graphs, stream-ordered allocation) to concrete LLM bottlenecks, helping engineers know which parts of the guide to apply.
- Wafer's thread frames speed work as finding where execution waits: an kernel waiting on its next data tile, or a serving runtime leaving the GPU idle until the next launch.
- Per the thread, the CUDA guide's async copy, TMA and producer/consumer pipeline sections explain how supported GPUs load the next tile while computing the current one, and what synchronization is needed to reuse buffers safely.
- The guide's virtual memory chapter maps directly to growing KV caches: reserve an address range, then map more physical memory into it as the cache grows, without relocating existing data.
- CUDA Graphs cut repeated CPU launch setup and stream-ordered allocation avoids broad synchronization around temporary memory, which matters when launch and allocation costs show up in request latency.
- The thread flags multi-GPU interference: a compute kernel's completion fence can wait on unrelated NCCL writes, and memory synchronization domains can reduce that on supported hardware.
- LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
- inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
- attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
postEtalon: measuring LLM streaming latency beyond TPOT averages
postInference scaling checklist drawn from the PaLM TPU v4 paper
postwe launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the seri
postWafer says it beat Cerebras on latency for YC's AI Office Hours
Checking sign-in…
Loading comments…
