ProcessKernel: Calibrated Process Supervision for GPU Kernel Generation
Abstract
Generating efficient GPU kernels requires searching a vast space of functionally equivalent programs whose performance depends on low-level hardware behavior. Effective search therefore relies on execution feedback to identify promising implementations. In practice, however, obtaining such feedback is cumbersome and expensive: candidate kernels must be compiled, validated, and evaluated under a reliable testing setup, making repeated execution a major bottleneck to scalable kernel search and generation. We thus introduce an execution-free process supervision framework that turns one-time outcome measurements into reusable guidance for unfinished kernels during autoregressive generation, enabling on-the-fly search without executing intermediate candidates. Extensive experiments show that our approach reduces process-supervision labeling cost by (9.3 vs. an estimated 707 H100 GPU-hours), while improving generated-kernel correctness by up to 38% and achieving up to 21.64 speedup. These results highlight the promising potential of reusable process-level supervision to amortize costly execution feedback and enable more scalable search for correct and efficient GPU kernels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.