KernelAudit: Can LLM-Generated GPU Kernels Be Both Reliable and Competitive on Real-World Workloads?
Abstract
Existing evaluations of LLM-generated GPU kernels report speedups over PyTorch and accept a kernel once it passes a functional check with a fixed error tolerance. Kernels that exploit the measurement or ignore real-world requirements still pass this check, so the reported speedups overstate what the kernels deliver. We present KernelAudit, a benchmark that asks whether kernels generated by frontier LLMs can be both reliable and competitive on real-world workloads. First, after correctness and validity audits, we construct 230 tasks across 14 domains from expert-written, community-maintained kernels collected from 69 high-star open-source repositories. Each expert kernel serves as both the speed baseline and the reliability baseline. Second, from an analysis of 505 generated kernels that passed functional checks, we build a verifier that audits every kernel at five levels, ordered by how strong a condition is needed to expose a defect. Third, we build a sealed evaluation environment that withholds all evaluation assets from the agent and times generated and expert kernels under one protocol on an exclusive GPU. Across seven frontier LLMs on all 230 tasks, only 25.3% of the generated kernels pass all five audit levels, and no model exceeds 40.9%. Giving the agent a fast audit tool built from the verifier during generation eliminates measurement bypass and algorithmic mismatch, raises the fraction of kernels that pass all five levels from 42% under an empty audit to 80%, and leaves speed unchanged. Agents optimize what is measured, and KernelAudit makes reliability measurable. Code and data are available at https://anonymous.4open.science/r/kernelaudit-BD8D.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.