← All IntelIntel / post
Big model. 2.4T params, 95B active. 4.89TB. Surprised they didn't release a 4-bi
- Source
- x.com
- Author
- alexocheema
- Date
Why it matters
A 2.4T-parameter open release is not locally runnable as shipped: 4.89TB of weights with no 4-bit quant from the lab, and even 4-bit needs roughly four 512GB machines.
Terms in this piece · Glossary
- inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Read the source x.com
More from alexocheema
- postIs the Rush to Build New LLM Inference Engines Fragmenting the Ecosystem?
post@Jason @Lons @eisokant @ape Correction: Should actually be closer to 50 tok/sec
postApple markets a four-Mac cluster for trillion-parameter local inference
postLocal inference crosses over: the M5 Ultra and the 1,000x claim
Recommended reads
Comments
Checking sign-in…
Loading comments…

