acceptodds
Under review as a conference paper at ICLR 2027

CUDACoach: Guiding CUDA Kernel Optimization with Diverse Strategies Learned from Verified Code

Abstract

Optimizing CUDA kernels is critical for high-performance deep learning systems. Recently, multi-role iterative frameworks powered by large language models have advanced this optimization process by establishing an automated “planning-generation-execution” pipeline, which effectively enhances kernel correctness and performance. However, improving strategy planning within these frameworks requires addressing two challenges: (a) Off-the-shelf planners repeat generic strategies across kernels with different performance profiles, limiting the search for faster implementations. (b) Supervised training to improve these recommendations requires reliable strategy demonstrations, but successful kernel implementations alone do not provide explicit modification guidance. We introduce CUDACoach, a framework for learning repair and optimization strategies to guide a frozen coder. (a) To obtain reliable demonstrations for supervised initialization, we propose verified strategy mining, which reconstructs strategies from successful kernel edits and validates them through implementation by the code model. (b) To counter repetitive optimization strategies, we use a two-level diversity reward during reinforcement learning to encourage distinct technique combinations among candidates and downweight techniques frequently used across training. We evaluate CUDACoach on all 250 KernelBench tasks with three downstream coders. With Qwen3-Coder-Plus as the coder, our method achieves 60.4% correctness and produces correct kernels faster than PyTorch on 19.2% of tasks, compared with 44.8% and 13.2% for single-stage iterative refinement and 42.4% and 13.6% when the coder also serves as the planner. Ablations confirm the benefits of diversity-reward training and the multi-kernel context for kernel correctness and optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.