CORAL: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization
Abstract
GPU kernel optimization requires generating correct implementations, identifying effective optimization techniques, and applying them to suitable workloads. We present CORAL, a reinforcement learning framework that couples skill discovery with skill exploitation within a single shared language model through a dynami- cally evolving skill library. CORAL jointly trains three agents sharing one LLM backbone: a Skill Selection Agent that retrieves relevant techniques via BM25 and LLM reranking, a Policy Agent that generates multi-turn Triton kernels conditioned on selected skills, and a Skill Summary Agent that distills successful rollouts into reusable skills. Candidate skills are added only after execution-based verification confirms first-turn speedups on their source tasks. The agents are initialized via a structured SFT cold start on diversity-filtered data and jointly optimized with multi-turn REINFORCE and per-agent advantage estimation. On KernelBench, CORAL-14B achieves 37.2%, 70.6%, and 32.2% on Level 1, 2, and 3 under the Fast threshold using the last output of three-turn rollouts, outperforming our repro- duced Dr. Kernel-14B baseline. Ablation studies show benefits from inference-time skills, joint training, learned selection, and diversity filtering, with trade-offs across task levels and speedup thresholds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.