Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading
Source
developer.nvidia.com
Author
Tanya Lenz
Date
Why it matters
When HBM binds before compute does, offloading activations to host memory can beat recomputing them by up to 57% — and unlock batch sizes that were previously impossible — provided the transfers actually overlap.
Terms in this piece · Glossary
mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.