acceptodds
Under review as a conference paper at ICLR 2027

GPU Forecasters: Language Models as Speculative Surrogates for Kernel Runtime Optimization

Abstract

GPU kernels are the workhorse of modern deep learning, and optimizing them (via evolutionary search or coding agents) usually requires repeated measurement on target hardware. While these measurements provide the ground-truth signal necessary for kernel search, they are costly, because each evaluation of a kernel requires compilation and repeated execution on a GPU. As improvements in LLM inference reduce the cost of writing novel kernels and LLM-driven searches scale to large search budgets, on-device evaluation becomes a bottleneck. To address this, we study how LLMs can serve as calibrated, speculative surrogates of the GPU: probabilistic performance models that forecast the speedup of a candidate kernel from its source. During a kernel search algorithm’s generation phase, the surrogate forecasts runtimes for generated kernel, and only the top-ranked kernels are scored by the GPU. The search uses the speculative forecasts by the surrogate as scores for the remaining kernels. We first evaluate four off-the-shelf LLMs as surrogates. We find that LLMs can act as effective surrogates, helping searches find fast kernels while measuring only a fraction of generated candidates on the GPU. On 1,347 parent-child pairs from real searches, LLMs identify discovery moments, i.e., children that are much faster than their parents, at rates above chance. We then integrate a surrogate into evolutionary kernel search to prioritize candidates for GPU measurement. Measuring only the most promising generated candidates, this search matches or outperforms GPU-only search on most benchmark tasks. Finally, we train a surrogate with reinforcement learning and find that calibration rewards improve calibration without hurting ranking, while correctness rewards alone make confidence less reliable. We release the 12,388 measured LLM-generated kernels produced by these searches. These results suggest that LLMs can play a broader role in kernel optimization, by acting as virtual models of a GPU rather than solely as kernel generators for search.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.