
DeepGEMM
github.com/deepseek-ai/deepgemm- Category
- Developer Tools
- Pricing
- Open Source
- Type
- TOOL
- Builder
- deepseek-ai
- GitHub
- 7.6k stars
- Added
- May 4, 2026
About
A high-performance CUDA library for FP8/FP4 tensor operations in large language models, featuring optimized GEMM kernels, MoE fusion, and runtime JIT compilation. Designed for NVIDIA GPUs with clean, accessible code for learning GPU optimization techniques.
Why it made the leaderboard
DeepGEMM represents serious GPU optimization infrastructure from DeepSeek that addresses critical performance bottlenecks in LLM inference and training. The combination of FP8/FP4 precision, MoE fusion, and JIT compilation shows genuine technical depth, and while niche, it's exactly the kind of foundational tooling that enables efficient AI model deployment at scale.
Tags
cudagpumachine-learningtensoroptimizationllmnvidiaperformance
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.