Vibeleaderboard
← All Intel
Intel / article

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Source
arxiv.org
Author
William Fedus et al.
Date
Why it matters

The paper that made practical, by routing each to exactly one expert instead of blending many. Every frontier model advertising a huge parameter count with a small active count runs this idea — which is why parameter count and cost stopped being the same conversation.

Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Recommended reads
Comments

Checking sign-in…

Loading comments…