
Explains how Moonshot AI keeps a 2.8T-parameter model deployable via LatentMoE compression, a KV-cache-capping hybrid scheme, and aggressive down to community 1-2 bit GGUF — concrete architectural patterns engineers building large systems can study and reuse.
“Kimi K3 is a multimodal Mixture-of-Experts (MoE) model with 2.8 trillion total parameters. Through extreme sparsity, only about 104 billion parameters activate per token, roughly 16 out of 896 experts. It natively supports a 1 million token context window.”
Ben Dickson
“Moonshot AI recently released the weights for its flagship Kimi K3 model, making it the largest open-weights LLM to date, rivaling Opus 4.8 and GPT-5.6.”
Ben Dickson
“on the AA-Briefcase benchmark it places second overall, trailing only Claude Fable 5, and beats both GPT-5.6 Sol and Opus 4.8 in that category”
Ben Dickson
“You cannot spin up 2.8 trillion parameters efficiently without innovations at every level of the stack.”
Ben Dickson
postPrompt cache TTL is the hidden line item in long coding-agent sessions
articleFull breakdown: Claude-only apps often wait until the host speaks the new revisi
articleDo not ask whether the agent follows the rule. Ask what stops it when it does no
articleFull Breakdown: When can a change ship unread? When something cheap and hard to Checking sign-in…
Loading comments…