KernelCompass: Traversing the Feasible Resource Space of the Target GPU Toward a Task-Contracted Objective
Abstract
GPU kernel optimization must account for both the hardware and the user's goal: a kernel that is faster in isolation need not make the model faster under its intended workload. We present KernelCompass, an agent framework that combines hardware resource limits, measured bandwidth and compute throughput, and tuning records to adapt kernel resource demand to the target device. Within one algorithm structure, changing a single parameter while holding the others fixed yields paired performance and resource feedback that guides rewrites to relieve resource constraints while using available headroom. Rewritten structures are retuned, and the new measurements guide the next round. At the model level, a task contract fixes the metric, workload, quality conditions, and acceptance criterion before optimization starts, and drives both the choice of operator groups and the acceptance of candidates under the native workload. With GLM-5.3-max as the base model, KernelCompass is to faster than the strongest of four baseline frameworks in each of the 18 combinations of six KernelBench tasks and three GPUs. On Qwen3-4B on an RTX 4090, separate goal-specific optimizations reduce time to first token (TTFT) by 3.56%, raise single-stream throughput by 12.51%, and raise concurrent throughput by 9.29% under the respective tested workloads.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.