← All IntelClip / OtherRoadmap for weight compression: FP4, KV cache, sparsity
From Compression at the Edge — Chris Alexiuk, NVIDIA · ≈39:49
“quantization does not degrade accuracy that much but spec uh sparity causes accuracy degradation a bit more.”
“deepseek started it I blame them they started with MLA”
“so we're going to quantize more things uh more is basically the idea”
What’s in it
- Explains why weight quantization is nearing its practical limit
- Previews Nvidia Rubin's new dynamic activation sparsity feature
- Breaks down why sparsity trails quantization in real-world adoption
Clip transcript
yeah, >> going to phones. Let's go. >> Yeah. Okay. So the way I think about compression, uh the whole space is going to go more broader. So we will focus this discussion mostly on weight compression. So uh okay, going back to FP4 once again. Uh so it also do weight compression and uh math acceleration because it it do the uh gem in 4bit. So uh yeah so for in terms of weight compression we might be able to go to like one maybe two or three bit more but in terms of u quantization alone we might be like close like close to the par optimality uh yeah let's let's see uh then there are more like you know more type of compressions so KB cache compression right so yeah so people are still mostly using 8 bit uh 4-bit large models retain that quality very well but smaller models we see some drop. So yeah so looking forward to KV cache uh maybe comb action plus quantization uh like pushing uh the long uh horizon reasoning uh broader than sparsity. So uh so okay so by the way sparsity is uh has been part of Nvidia hardware but it has not been uh like broadly adopted. uh that is because um quantization does not degrade accuracy that much but spec uh sparity causes accuracy degradation a bit more. So in Rubin there is this cool feature called dynamic activation sparity so it can uh improve attention math etc. So yeah so looking forward to that and then coming back to like you know all these heterogeneous architectures right so with each release the yeah the particularly the attention architecture is getting more and more complex deepseek started it I blame them they started with MLA now like you know sparse attention indexed attention like you know lot of skips soft max yeah so it's getting broader and yeah so it is part of this process where we make models cheaper and and they are they are having compounding effects Perfect. So, so we're going to quantize more things uh more is basically the idea. That's that's pretty dope. I I don't I don't mind that. Take us home.
Comments
Checking sign-in…
Loading comments…