acceptodds
Under review as a conference paper at ICLR 2027

Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for optimized code generation. However, existing RLVR approaches primarily rely on outcome-based signals such as correctness and speedup, overlooking performance-critical structural properties of programs that are essential for generating optimized code. In this work, we propose \ourtool, a reflective RL framework that incorporates both verifiable execution rewards and structural code-aware rewards derived from parallelization features (e.g., memory coalescing, occupancy, Arithmatic Intensity, and synchronization patterns). \ourtool operates in two stages: (1) an offline pairwise ranking module that learns to distinguish strong and weak program candidates via contrastive comparisons, and (2) an online RL training phase that jointly optimizes for correctness, performance, and structural efficiency through a unified reward signal. To further enhance learning, \ourtool utilizes iterative refinement using execution feedback enabling progressive improvement of generated candidates. We also introduce a dataset comprising 2.9k C CUDA and 1k PyTorch CUDA programs, each paired with diverse input configurations and multiple CUDA implementations encompassing diverse optimization strategies. \ourtool is evaluated across multiple benchmarks comprising both C CUDA and PyTorch CUDA transformations. Empirical findings suggest that \ourtool significantly outperforms strong baselines, including Qwen-3-32B (for C CUDA) and CUDA Agent (for PyTorch CUDA) by achieving up to 5X & 3.32X improvements in speedup, and 17% & 7% improvements in correctness, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.