REPAIR: Learning What to Optimize for Efficient Ascend Kernels
Abstract
High-performance kernels are essential for fully exploiting modern accelerators. Large language models (LLMs) offer a promising route to automated kernel generation and optimization, but struggle with the specialized requirements of Ascend kernels. Existing end-to-end approaches, including reinforcement-learning-based and agent-driven optimization, yield limited performance gains in this setting. We argue that a key challenge lies in coupling optimization decision-making with code implementation, requiring models to jointly determine what to optimize and how to implement the corresponding changes. To address this challenge, we propose REcursive kernel oPtimizAtion by Identification and Revision (REPAIR), a recursive optimization framework that explicitly decouples these two responsibilities. At each iteration, Identification selects an actionable optimization direction without modifying the kernel, while Revision takes this decision as input and implements the corresponding code changes. Building on this separation, we leverage publicly available agents to generate data for both stages and train a compact LLM with reinforcement learning specifically for Identification. This design focuses RL on learning which optimization to pursue, rather than jointly learning optimization decisions and complex code modifications, allowing the compact model to specialize in decision-making while delegating implementation to Revision. Experiments on 253 distinct Ascend kernels from MultiKernelBench show that REPAIR achieves up to 34.8× speedup over the original kernel code, outperforming the 9.5x for the baseline Codex agent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.