- Category
- Developer Tools
- Pricing
- Open Source
- Type
- TOOL
- GitHub
- 1 stars
- Date
About
Kairo is a research workbench for LLM inference on NVIDIA Blackwell GPUs that measures specific workloads (such as CUDA Graph replay versus eager execution for NVFP4 serving on an RTX 5090) and only promotes exact, measured workload configurations into a fail-closed runtime routing policy. It is explicitly not a general-purpose inference engine and does not replace vLLM, CUTLASS, or cuBLAS.
Why it made the leaderboard
It gives LLM infra engineers a reproducible, evidence-gated methodology for deciding when CUDA Graph replay actually helps NVFP4 inference throughput instead of guessing.
Tags
inferencegpublackwellcudabenchmarkingllm
Tech Stack
Python
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.
