KernelMaster: Agentic GPU Kernel Optimization from Semantics to Execution
Abstract
Agent systems can generate and refine GPU kernels, yet optimizing an unfamiliar operation often requires extensive trial and error. GPU kernel optimization requires coordinated choices about how to reorganize computation, distribute and schedule parallel work, and implement it using GPU instructions. However, runtime and profiler feedback reveal performance bottlenecks but do not directly prescribe effective program changes, making it difficult for an Agent to connect these observations to actionable revisions without repeated trial and error. We present KernelMaster, a system that guides Agent-driven GPU kernel optimization through mathematical dataflow search and hardware-calibrated execution synthesis. We expose mathematical transformations through an Agent-operable MathIR, validate candidates through source coverage and replay, and project selected dataflows into LoopIR. We then synthesize Warp pipelines and CTA programs using target-GPU measurements of service time, concurrency, and HBM bandwidth. The selected execution graph guides Agent-driven CUDA implementation and refinement. On five newly specified attention kernels, KernelMaster achieves – speedup over the strongest Agent baselines and – over PyTorch references. Across established kernels on Hopper and Blackwell, it approaches or exceeds expert implementations, with speedups up to .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.