AscendOptimizer: Episodic Agent for Ascend NPU Operator Optimization
Abstract
Optimizing AscendC (Ascend C) operators for Ascend NPUs is difficult for two reasons. First, unlike CUDA, the ecosystem offers few public kernels to learn from. Second, performance depends on a coupled two-part implementation: a host-side tiling program that controls data movement and a kernel program that schedules and pipelines computation. We present AscendOptimizer, an episodic agent that builds missing optimization knowledge from execution itself. For kernel optimization, AscendOptimizer rewinds strong implementations by removing optimizations in a controlled way, then keeps the changes whose removal measurably hurts performance as reusable experience for later rewriting. For host-side optimization, it runs profiling-in-the-loop evolutionary search to find valid, fast tiling and data-movement configurations directly from hardware feedback. On a transductive benchmark of 101 real AscendC operators, AscendOptimizer achieves a 1.21× geometric-mean speedup over the open-source references, and 53.47% of operators run faster than those references. With hardware-evaluation counts matched across methods, its aggregate geometric-mean speedup exceeds Best-of-N sampling, OpenEvolve, CudaForge, and EvoKernel.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.