🚀 Meet HySparse: Our new breakthrough in long-context LLM efficiency! We’re excited to share HySparse (Hybrid Sparse Attention)—a hybrid model architecture that interleaves each full attention layer with multiple sparse attention layers, where the sparse layers strategically derive important token selection and KV caches from the preceding full layer! 📖 Paper link:

🔍 What makes HySparse different? HySparse with Oracle Token Selection and KV Cache Sharing resolves two fundamental limitations of prior sparse attention methods: ⚠️ Inaccurate Proxy Selection: Conventional methods rely on approximate "proxies" or additional modules to guess token importance. HySparse uses the full attention layer as a precise oracle to identify important tokens. No more proxy guessing which tokens matter. 💾 High Memory Overhead: Existing dynamic sparse attention approaches reduce computation but fails to reduce KV cache storage. Instead of creating new memory for every layer, HySparse enables sparse attention layers to reuse the full attention KV cache.
📈 Proven Performance HySparse demonstrates in experiments that an 80B MoE model requires only 5 out of 49 layers to utilize full attention, enabling a nearly 10x reduction in KV cache storage. Despite this efficiency, the model improves performance over full attention baselines on general and long-context benchmarks.
Sparse usually cuts compute but not memory. Deriving token selection and from a preceding full-attention layer attacks the KV cache itself, which is the binding constraint on long- serving.
postHeads up, agent users! If you're using Xiaomi MiMo with thinking mode: When…
postMiMo-V2.5 and V2.5-Pro go open weights under MIT with day-zero SGLang and vLLM
postIntroducing MiMo-V2.5 Voice — our full-stack voice lineup for the Agent era. 🚀…Checking sign-in…
Loading comments…