mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
Why it matters
Mistral 3 puts a 675B-total / 41B-active mixture-of-expertsA model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.Full definition → flagship under Apache 2.0 alongside 14B, 8B, and 3B dense models, with image understanding and strong non-English multilingual performance. The weights can be self-hosted.