acceptodds
Under review as a conference paper at ICLR 2027

Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization

Abstract

Recent advances in large language models (LLMs) have enabled automated kernel generation and optimization, but most existing approaches rely on surface signals such as compilation feedback and profiling metrics, which reveal that a kernel is slow but not why the backend compiler fails to realize a profitable optimization, especially on emerging accelerators such as NPUs. We therefore formulate kernel optimization as a progressive cross-layer diagnosis problem that links runtime symptoms to IR structure and compiler behavior before rewriting source, and present a compiler-grounded, hierarchical optimization framework for Triton kernels that escalates from lightweight pattern triage and profiling diagnosis to IR attribution and compiler-grounded analysis only when deeper evidence is needed. Implemented on Triton for Ascend NPUs and evaluated on 37 successfully converted entries from a standardized NPUKernelBench-derived Ascend 950 benchmark, the system attains a geometric-mean speedup of and a median of from the initial to optimized Triton kernel; 22/37 exceed and 13/37 exceed . The complete distribution ranges from near-baseline entries to large wins, motivating transparent reporting of the current system's scope and limitations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.