If you're evaluating whether to trust LLMs to generate performance-critical GPU code, ParallelKernelBench gives you a grounded reality check: frontier models solve fewer than a third of multi-GPU kernel workloads, yet occasionally produce kernels that outperform any public implementation.
ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads.
The best model solves under a third, but a few generated kernels beat any public implementation.
Transcript
ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.