SparseKernelBench: A Benchmark for Sparse GPU Computation
Abstract
Sparse GPU performance depends on both the number and locations of nonzeros, requiring evaluation across input structures. We introduce SparseKernelBench, with 41 tasks across seven sparse representations, 533 performance cases and 637 correctness-only cases. Each task specifies input/output requirements and input families that a program must support. Four large language models generate programs without execution feedback for every task. Haiku produces no custom kernels. Programs from the other three models outperform the fastest tested library baseline, including compiled framework compositions, on 36.6–54.2% of the 533 performance cases. Requiring every task check to pass gives 36.0–44.3%, with the same denominator. Three of 103 programs that pass every performance-input check fail dedicated correctness-only inputs. At fixed dimensions, nonzero count and feature width, the fastest generated sampled dense–dense multiplication program takes 1.98× the fastest tested baseline’s runtime on uniform inputs but 0.54× on skewed inputs. Changing the baseline formulation also reverses whether programs outperform the library on 29–64 cases per model. These results show why sparse evaluation must establish the input domain, numerical requirements and tested comparator behind a performance claim.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.