KernelGenBench: Can LLMs and Agents Write Efficient Kernels Across Operator Sources and Hardware Platforms?
Abstract
Modern AI systems depend on specialized accelerator kernels, whose development is complicated by increasingly diverse operators and hardware. LLMs and agentic systems promise to automate this work, but existing evaluations do not show how performance varies across operator sources and hardware platforms, or what that variation costs. We present KernelGenBench, the first unified multi-source × multi-chip infrastructure for evaluating LLM- and agent-generated Triton kernels. Unlike benchmarks that require platform-specific target languages, KernelGenBench uses Triton as a common programming target across six hardware platforms. We report two controlled analytical views: KernelGenBench-MS (Multi-Source) covers 210 operators from PyTorch ATen, production vLLM operators, and proprietary cuBLAS routines, while KernelGenBench-MC (Multi-Chip) evaluates a semantically stable 110-operator subset across six hardware platforms. Our evaluation consumed over 15 billion tokens. More interactive agentic workflows achieved higher overall correctness, but no method dominated across sources and platforms: vLLM posed the strongest correctness challenge, cuBLAS set the highest performance ceiling, and AutoKernel with GLM-5.0 fell from 87% on NVIDIA to 25% on Iluvatar CoreX. Higher correctness came at substantial cost: specialized agents averaged 4.99 million tokens per successful operator, rising to 6.25 million for CUDA Optimized Skill. The results establish operator source, hardware platform, and generation workflow as distinct dimensions of kernel-generation capability, and show that success in a familiar source–hardware setting is not a reliable proxy for deployment readiness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.