QKernelBench: Benchmarking LLMs for Quantized GPU Kernel Generation
Abstract
Quantized GPU kernels are central to efficient LLM inference, yet kernel benchmarks judge them as they judge dense kernels: by comparing a candidate's outputs with a reference's under a tolerance. A quantizer that computes its own scale returns quantized values (codes) together with that scale, and the codes must be judged together with it: the same codes can be correct under a scale one ULP from the reference's and wrong under the reference's own. A check that compares codes and scale with the reference's separately rejects correct kernels, accepts wrong ones, or both. We introduce QKernelBench, a benchmark of 84 CUDA kernel contracts for quantized inference. Each contract states its numerical semantics clause by clause, and its verdict follows from them: codes are accepted when the returned scale is admissible and a declared arithmetic path reproduces them bit for bit, other codes must be exact outside declared one-step masks, and floating-point outputs must lie within declared bounds: a number of representable steps, a bound derived from the declared accumulation, or a tolerance. Every verdict is calibrated against mutants of its own contract and reports the positions it rejects. We judge kernels written by eight LLMs in two rounds, the second generated after the contracts were frozen. Frontier models solve 71 to 80 of the 84 contracts in the first round, but their kernels reach only a fraction of the speed of the shipping kernels we timed. On the same kernels, the correctness rules of published benchmarks err in both directions: they reject kernels whose scale differs lawfully from the reference's and accept kernels that break a stated clause, and no distance threshold separates the two. Under KernelBench's allclose, each frontier model would be credited with 14 to 18 fewer contracts. After repairs that the shipping kernels of eight production libraries prompted, the contracts accept their scale paths, reassociation and tensor-core accumulation; the kernels still rejected follow another convention for a stated clause, such as rounding, saturation or the zero guard, compute in a lower precision than the contract admits, or return codes that change between calls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.