Virtuoso Lite Virtuoso Medium V2 Distilling Deepseek V3 Into 10b 32b Small Language Models Slms
Source
Arcee AI editorial sitemap
Author
Arcee AI editorial sitemap
Date
Terms in this piece · Glossary
distillation — Training a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Why it matters
Logit-level distillationTraining a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.Full definition → across mismatched tokenizers is unusual, and it produced a permissively licensed 10B that holds its own against a 32B, which changes what fits on modest hardware.