MegaKernelBench: A System-Level Benchmark for Megakernel Generation and Optimization
Abstract
Small-model, small-batch LLM decoding launches many fine-grained GPU kernels, making launch overhead a bottleneck for interactive inference. Megakernels mitigate this overhead by executing part or all of a computation graph within a single persistent work kernel. However, their performance depends jointly on operator implementation and scheduling, which existing GPU kernel benchmarks do not adequately evaluate. We present MegaKernelBench, a system-level benchmark for megakernel generation and optimization with four progressive tasks: operator generation, scheduling generation, joint operator-scheduling generation, and generation without a prebuilt runtime. It evaluates correctness, structural validity, and end-to-end performance. Existing methods can generate correct megakernels with a prebuilt runtime but generally underperform mainstream inference engines; generation without one remains more challenging. Strong operator-scheduling coupling further makes joint optimization essential. MegaKernelBench enables unified, reproducible evaluation of megakernel generation by compilers and agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.