A new Apache 2.0 open weightsA model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.Full definition →mixture-of-expertsA model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.Full definition → with 3.5B active parameters and 262k context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → is a self-hostable option for German and English workloads with data-residency needs. The post covers when it fits and where it falls short.
Terms in this piece · Glossary
open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.