How GLM5.3 Sparse Attention Affects HBM Memory Usage
Source
Kimbo Chen
Author
Kimbo Chen
Date
Key takeaways · AI-distilled
SemiAnalysis's InferenceX puts GLM-5.3 on GB200 at about $0.044 per million total tokens at 150 tokens/s versus $0.238 on MI355X with ATOM, but those GB200 runs had 14-19s p90 time to first token. Capped at 2s TTFT, B200 was still about 47% cheaper.
In B200 runs, raising concurrency from 8 to 16 requests cut prompt-token reuse from GPU memory from 90.3% to 54.8% while reuse from host memory rose from 6.0% to 40.3%, keeping the overall cache hit rate above 95%.
GLM-5 is a 744B-total, 40B-active mixture-of-expertsA model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.Full definition → using DeepSeek Sparse attentionThe mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.Full definition →. Its 64 query heads, half of DeepSeek V3.2's, give about 120.8 FLOP/B arithmetic intensity, which SemiAnalysis suspects targets Moore Threads' MTT S4000 rather than Nvidia's H800.
GLM-5.2's IndexShare lets every four DSA layers share one indexer, caching top-K indices so shared layers can reuse the selection. Z.ai reports 75% less indexer cache and indexer FLOPs and 1.5x to 1.8x higher throughput across context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → lengths.
For long-horizon RL, GLM-5.2's Single-rollout Asynchronous Optimization replaces GRPO's group-normalized advantage with value-model GAE that needs one rollout per prompt, easing straggler latency at the cost of extra compute and a doubled memory footprint.
Terms in this piece · Glossary
attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
Why it matters
Sparse attention cuts compute per token but top-k selection still needs the full KV cacheThe memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.Full definition → in HBM. Tiered designs like HiSparse trade cache-miss I/O for capacity, which affects long-context serving cost.