Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Pricing
Open Source
Type
TOOL
GitHub
1 stars
Date

About

Kairo is a research workbench for LLM inference on NVIDIA Blackwell GPUs that measures specific workloads (such as CUDA Graph replay versus eager execution for NVFP4 serving on an RTX 5090) and only promotes exact, measured workload configurations into a fail-closed runtime routing policy. It is explicitly not a general-purpose inference engine and does not replace vLLM, CUTLASS, or cuBLAS.

Why it made the leaderboard

It gives LLM infra engineers a reproducible, evidence-gated methodology for deciding when CUDA Graph replay actually helps NVFP4 inference throughput instead of guessing.

Tags

inferencegpublackwellcudabenchmarkingllm

Tech Stack

Python

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.