← All IntelClip / EntertainmentGPU memory hierarchy: SRAM vs HBM
From Why LLM Inference Is Memory-Bound (Julia Turc) · ≈1:43
“NVIDIA's H100 GPU has a bit over 50 megabytes of SRAM.”
“An H100 comes with 80 gigabytes, much more suitable for storing the model weights.”
What’s in it
- Explains the GPU memory hierarchy: SRAM vs HBM tradeoffs
- Learn why H100's 50MB SRAM can't hold large LLMs
- Understand where model weights actually live during inference
Clip transcript
So the weights need to be stored somewhere. GPUs typically come with their own on-chip memory. This is static RAM or SRAM. It's fast but expensive to manufacture and takes a lot of physical space, which is why there's only so much of it. NVIDIA's H100 GPU has a bit over 50 megabytes of SRAM. That's not nearly enough for state-of-the-art LLMs. Open weight models go up to hundreds of gigabytes of storage and proprietary ones must be even larger. To support this demand, GPUs come with yet another level of memory placed near the processor, but not on the same chip. It's called high bandwidth memory or HBM. This is typically what people refer to when they say "GPU memory". Unlike SRAM, it uses a stacked die architecture that trades some speed for much greater capacity. An H100 comes with 80 gigabytes, much more suitable for storing the model weights.
Comments
Sign in to comment.
Loading comments…