Vibeleaderboard
Index / tool

MiniMax Sparse Attention (MSA)

github.com/minimax-ai/msa
Visit github.com
Category
AI Tools
Rank
No. 2016Tools index
Pricing
Open Source
Type
TOOL
GitHub
419 stars
Date

About

MSA (the fmha_sm100 package) provides dense FlashAttention and sparse top-k attention kernels for NVIDIA SM100 (Blackwell) GPUs. This Fireworks integration branch vendors a KV-outer sparse prefill backend developed by Fireworks AI and routes sparse prefill through it when GQA is at least 8, falling back to the original MiniMax MSA CuTe kernels below that, all behind an unchanged fmha_sm100 / fmha_sm100_plan Python API. The practical effect is faster prefill for MiniMax M3 inference without callers changing their code.

Tags

attentioncudainferenceflashattentionblackwellminimax

Tech Stack

Python

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.