Efficient Memory Management for Large Language Model Serving with PagedAttention
Source
arxiv.org
Author
Woosuk Kwon et al.
Date
Why it matters
The reason an inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → server can hold many concurrent requests: treating the KV cacheThe memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.Full definition → like virtual memory pages instead of one contiguous block removes the fragmentation that was wasting most of the GPU. It is the paper behind vLLM, which is what most self-hosted serving actually runs.
Terms in this piece · Glossary
KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.