KernelWeaver: A Shape-Aware Kernel Agent for GPU Kernel Optimization
Abstract
Large language model (LLM)-based kernel agents automate GPU kernel development, but the kernels they generate are often optimized for only a few representative input shapes and struggle to maintain performance across the broader range encountered at deployment. This mismatch arises because shape changes alter the dominant hardware bottleneck, requiring different execution configurations. Although generating a dedicated kernel for every shape can address this mismatch, it increases compilation and code-size costs. We present KernelWeaver, a shape-aware kernel agent that optimizes cross-shape performance by composing a small set of complementary kernels under deployment budgets. To construct this portfolio, KernelWeaver first partitions the shape space using hardware resource constraints to guide candidate generation, then selects kernels according to their marginal performance gain per unit cost. The resulting portfolio is deployed through a shape-aware router, whose execution feedback identifies gaps in offline coverage and guides targeted refinement. On nine common GPU kernel workloads, KernelWeaver achieves geometric-mean speedups over PyTorch eager of on the NVIDIA RTX 4090 and on the NVIDIA A800, compared with and for the strongest evaluated baseline on the respective platforms. In our deployment cost evaluation, KernelWeaver reduces deployment compilation time by 91.71% compared with torch.compile.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.