AscendAgent: Evolving Context to Fit the Harness in CANN Kernel Optimization
Abstract
High-performance accelerator kernels are fundamental to efficient large language model (LLM) inference, yet optimizing them still requires substantial systems expertise. Existing approaches often optimize workloads independently, making it difficult to reuse experience across input patterns and hardware. We present \ascendagent, a kernel optimization framework that evolves reusable knowledge across workloads while keeping the optimization harness fixed. It organizes API documentation, community implementations, and technical discussions into conditional, composable transformations with expected effects and supporting evidence. \ascendagent operates through two mechanisms. (1) Within each workload, a frozen knowledge snapshot guides evolutionary search over a program graph. The harness proposes changes, checks correctness, and measures latency. Compatible programs share evaluation evidence and explored continuations, while LLM-guided mutations explore changes beyond existing knowledge. (2) After each source search, workflow replay uses recorded proposals and execution outcomes to revise transformation preferences, application conditions, and reusable combinations. Independent confirmation determines which updates guide later workloads. Model weights, the workflow, and the evaluator remain fixed, and each new workload starts with fresh local search state. We evaluate \ascendagent with five language models on four kernel benchmarks spanning Ascend NPUs and NVIDIA GPUs under a budget of candidate compilations per task. With \MainModel, \ascendagent achieves mean best speedups of 1.40 on CANN-Bench and 1.75 on TritonBench-G, compared with 1.21 and 1.36 for the same harness with frozen initial knowledge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.