acceptodds
Under review as a conference paper at ICLR 2027

Structure Before Tuning: Two-Stage Search for LLM-Generated GPU Kernels

Abstract

GPU kernels form the computational foundation of modern artificial intelligence systems, and their execution efficiency directly affects training and inference latency, throughput, and deployment cost. Large language models can generate kernel code and iteratively optimize it using compilation and runtime feedback, offering a new way to reduce the cost of manual tuning. However, LLM-generated kernels often fall short of the desired performance: reaching a high-performance implementation typically consumes substantial model tokens and evaluation budget, and the resulting optimization is unstable. Existing methods commonly search over complete kernel programs through a unified generate–compile–benchmark–revise loop that alternates between structural decisions and local implementation tuning. Structural decisions determine how computation is decomposed and coordinated, whereas local optimization refines implementation details within a given structure. When the search moves to a new structure, local refinements made for the previous structure often cannot be reused, reducing the value of the budget already spent and potentially causing candidates to be compared or discarded before they are sufficiently optimized. We introduce , a two-stage framework that decouples structural search from local optimization. In the first stage, an agent selects and composes actions from a predefined structural action library to construct abstract structural blueprints, which are then instantiated as executable domain-specific-language (DSL) implementations to explore alternative computation organizations. After structural exploration, promising structures are selected based on the search outcomes, and each retained structure becomes an independent starting point for the second stage. For each structure, an agent independently performs local implementation search under its corresponding structural contract. By exploring structures first and then optimizing each retained structure within its own trajectory, avoids repeatedly switching between structural search and local tuning and enables different structures to be compared after a controlled degree of implementation refinement. On FlashInfer-Bench and Sol-ExecBench kernel workloads running on NVIDIA A800 and H100 GPUs, we compare with mixed search, structure-only search, and local-only optimization. Insert final experimental results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.