Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
- Source
- latent.space
- Author
- Latent Space
- Date

Kimi K3 offers claimed frontier-class agentic coding performance with a 1M-token at Sonnet-class pricing, but the real decision factors are its heavy self-hosting demands (64+ accelerators), slower serving, and a regression — so it's most compelling if you need long-horizon coding on and can absorb the infra cost.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
- multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
- hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
- open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
“Arena later added that K3 has a 76% pairwise win rate in Frontend Code Arena, versus 63% for Fable 5 and 58% for GPT-5.6 Sol”
“Artificial Analysis published an independent evaluation placing K3 at 57 on the AA Intelligence Index , calling it comparable to Opus 4.8 and GPT-5.5 , but still behind Fable 5 and GPT-5.6 Sol overall”
“K3 uses Kimi Delta Attention (KDA) , which Moonshot says enables up to 6.3x faster decoding in million-token contexts”
“The strongest repeated theme: frontier-ish performance at materially lower price than top closed models , though not at bargain-basement open-model prices”
Checking sign-in…
Loading comments…





