← All IntelClip / AI ToolsQuantization degradation benchmarks
From Meta Open Source Is BACK – Muse Glimmer First Test! · ≈30:23
“So the trade-off here in degradation versus the trade-off in VRAM reduction to be able to run it is definitely worth it to test the smaller”
“the percent degragation right here that they measured for agentic tasks of this one that we're running today was only 0.2%.”
What’s in it
- Tests a model at ~4-bit precision instead of full precision
- Shows agentic-task degradation is just 0.2% at that compression
- Reveals shrinking further to fit 24GB VRAM only costs 1%
Clip transcript
closing thing, the final version that we were running here, if we go to the actual main page for this model, so the model is approximately at 4bit precision, so it's compressed to use this. Now, I could have tried this at full precision on a 6000 Pro Blackwell card, but I want to test it on something that more folks are likely going to be using. So, either one of these. And we can see the percent degragation right here that they measured for agentic tasks of this one that we're running today was only 0.2%. And then if you want to fit this in 24 gigs, it only drops to 1%. So the trade-off here in degradation versus the trade-off in VRAM reduction to be able to run it is definitely worth it to test the smaller
Comments
Sign in to comment.
Loading comments…