acceptodds
Under review as a conference paper at ICLR 2027

EvokeAgent: A Hardware-Aware, Knowledge-Base-Co-Driven Self-Evolving Agent System for GPU Kernel Optimization

Abstract

Automated GPU operator optimization is critical for accelerating deep learning workloads, prompting the growing use of Large Language Models (LLMs) to generate and tune high-performance kernels. While existing systems increasingly incorporate runtime profiling feedback, the selection of hardware evidence is often not explicitly coupled to optimization actions; relying on the model alone to analyze all such hardware data makes it difficult to pinpoint the right optimization opportunities amid noisy metrics. We present EvokeAgent, an NCU-native GPU operator optimization agent built around a self-evolving, bottleneck-indexed knowledge base (KB). EvokeAgent establishes a bottleneck-guided routing layer that contracts voluminous NCU metrics into hypothesis-relevant diagnostic evidence downward while mapping the diagnosis to targeted KB policy retrieval upward. Evaluated on all three KernelBench levels using an untuned LLM, EvokeAgent achieves 100% correctness with geometric-mean speedups of 1.93, 2.65, and 2.94 on CUDA, and 1.68, 2.56, and 1.64 on Triton across L1–L3, respectively. Furthermore, ablation shows that this routing-and-retrieval design outperforms the non-KB baseline by 1.08, 1.48, and 1.55 on L1–L3 under capped optimization rounds.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.