EvokeAgent: A Hardware-Aware, Knowledge-Base-Co-Driven Self-Evolving Agent System for GPU Kernel Optimization
Abstract
Automated GPU operator optimization is critical for accelerating deep learning workloads, prompting the growing use of Large Language Models (LLMs) to generate and tune high-performance kernels. While existing systems increasingly incorporate runtime profiling feedback, the selection of hardware evidence is often not explicitly coupled to optimization actions; relying on the model alone to analyze all such hardware data makes it difficult to pinpoint the right optimization opportunities amid noisy metrics. We present EvokeAgent, an NCU-native GPU operator optimization agent built around a self-evolving, bottleneck-indexed knowledge base (KB). EvokeAgent establishes a bottleneck-guided routing layer that contracts voluminous NCU metrics into hypothesis-relevant diagnostic evidence downward while mapping the diagnosis to targeted KB policy retrieval upward. Evaluated on all three KernelBench levels using an untuned LLM, EvokeAgent achieves 100% correctness with geometric-mean speedups of 1.93, 2.65, and 2.94 on CUDA, and 1.68, 2.56, and 1.64 on Triton across L1–L3, respectively. Furthermore, ablation shows that this routing-and-retrieval design outperforms the non-KB baseline by 1.08, 1.48, and 1.55 on L1–L3 under capped optimization rounds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.