← All IntelIntel / articleCan I Run This Model?
- Source
- Aadi1277
- Author
- Aadi1277
- Date
Terms in this piece · Glossary
- KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
Why it matters
The gap between a model's weight size and what it actually needs on the card is where most local deployments fail, and this shows the arithmetic.
Comments
Checking sign-in…
Loading comments…