Glueball: Agentic Generation and Verification of Hardware-Aware Gluon Kernels
Abstract
Modern GPUs provide increasingly powerful features to improve performance, yet effectively harnessing them requires reasoning about complex execution behavior. Large Language Model (LLM) agents are exceptionally well suited to exploring this optimization space through iterative code generation, with pre-defined test-suites executed on bare metal providing correctness guarantees and performance feedback. However, numerical testing alone provides limited coverage of the input and execution space, making it possible for agents to overfit to tested inputs while missing correctness violations arising from subtle hardware-execution semantics. To address these limitations, we present Glueball, the first system to equip LLM agents with theorem-proving tools to reason about correctness during GPU kernel generation and optimization. Glueball targets Gluon, a tile-level Domain Specific Language (DSL) within the Triton ecosystem that provides direct, fine-grained access to GPU hardware beyond the abstractions exposed by Triton's higher-level programming model. Our system coordinates optimization across single and multi-kernel workloads, using end-to-end model performance, as well as theorem-prover guidance, as optimization feedback to evolve the kernels. Our results show improved kernel bug detection, while evaluations on real-world workloads demonstrate end-to-end performance competitive with already performant Triton baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.