
MiniMax Sparse Attention (MSA)
github.com/minimax-ai/msa- Category
- AI Tools
- Rank
- No. 2016Tools index
- Pricing
- Open Source
- Type
- TOOL
- Builder
- @MiniMax_AI
- GitHub
- 419 stars
- Date
About
MSA (the fmha_sm100 package) provides dense FlashAttention and sparse top-k attention kernels for NVIDIA SM100 (Blackwell) GPUs. This Fireworks integration branch vendors a KV-outer sparse prefill backend developed by Fireworks AI and routes sparse prefill through it when GQA is at least 8, falling back to the original MiniMax MSA CuTe kernels below that, all behind an unchanged fmha_sm100 / fmha_sm100_plan Python API. The practical effect is faster prefill for MiniMax M3 inference without callers changing their code.
Tags
attentioncudainferenceflashattentionblackwellminimax
Tech Stack
Python
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.