Practical rules for shrinking models for local or edge serving, including a non-obvious long- failure mode that only shows up after deployment.
“So it's not 86% dumber right if you compress it by 86% it doesn't become like you know terrible useless.”
Daniel Han
“So the way I think about is same cost more intelligence.”
“there is a super weights paper which shows that if you quantize one number just one of the entire model your model becomes 20% dumber”
Daniel Han
“the fact that we could actually fine-tune stuff on a toaster essentially that is the collab tier T4 for me that was like a big wow”
Merve Noyan
“you do not need to be, you know, controlled by some model labs and now you can do whatever you like on your computer, right?”
Daniel Han
videoThe Dark Arts of Skill Engineering — Paul Bakaus, Renaissance Geek
videoOperating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta
videoVertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave
videoAre LLM Performance Benchmarks Reliable? — Ashok Chandrasekar & Jason Kramberger, GoogleChecking sign-in…
Loading comments…