Vibeleaderboard
← All Intel
Intel / article

Efficient Memory Management for Large Language Model Serving with PagedAttention

Source
arxiv.org
Author
Woosuk Kwon et al.
Date
Why it matters

The reason an server can hold many concurrent requests: treating the like virtual memory pages instead of one contiguous block removes the fragmentation that was wasting most of the GPU. It is the paper behind vLLM, which is what most self-hosted serving actually runs.

Terms in this piece · Glossary
  • KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Recommended reads
Comments

Checking sign-in…

Loading comments…