
-generated GPU kernels for SQL operators hit 2.11x over the baseline on TPC-H/H100, and the paper isolates what drives it: kernel fusion, full-query specialization for stronger models, workload over hardware context.
articleKernelArc: A Multi-Agent Framework for GPU Kernel OptimizationJoyjit Kundu, Ben Stoffelen, Kaili Wang, Peter Vrancx, Ludovic Denoyer
articleRealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI FrameworksJinjun Huang, Zhongzhen Wen, Tongtong Xu, Meng Yan, Xin Xia, Zhongxin Liu
articleParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)Together AIChecking sign-in…
Loading comments…