FlashAttention and ThunderKittens are foundational to modern LLM inference and training performance, so understanding how this team writes GPU kernels gives engineers insight into where real throughput gains come from.
The team behind FlashAttention and ThunderKittens — how Together AI's kernel researchers close the gap between GPU hardware and production AI.
Transcript
The team behind FlashAttention and ThunderKittens — how Together AI's kernel researchers close the gap between GPU hardware and production AI.