Transcript
one matrix replaced the KV cache.
(the technique is 100% open source)
Kimi just dropped K3, an open model at frontier scale.
it leans on a new mechanism called delta attention that does not keep a growing KV cache. that is how it holds a million tokens of context without the https://t.co/zEDCWlmi20 https://t.co/JCdqjCKcKU