acceptodds
Under review as a conference paper at ICLR 2027

PTXBench: Interpretable Evaluation of LLMs for Architecture-Specific GPU Programming

Abstract

Reaching peak performance on modern GPUs requires specialized hardware-accelerated mechanisms. Writing kernels that use them is therefore a key capability for LLMs. Yet correctness and speedup alone cannot show whether a generated kernel uses the intended hardware mechanisms. We propose instruction-qualified speedup, which counts a kernel's speedup only if the kernel is correct and executes selected tensor-compute or data-movement instructions. Our framework, PTXBench, measures this metric while controlling four conditions: supplied architectural knowledge, execution feedback, harness capabilities, and generation budgets. Ablations show that supplied architectural knowledge and execution feedback improve instruction-qualified speedup. Using a shared multi-turn refinement harness, we evaluate five recent LLMs on one GEMM and four attention workloads across NVIDIA H100 and B200. Correctness is not enough. Even with full knowledge and execution feedback, 25.9% of trajectories that produce a correct kernel never produce an instruction-qualified one. In all ten GPU-workload settings, the fastest CUDA-PTX kernels execute both accelerated compute and data movement. Still, no model’s median instruction-qualified speedup reaches the vendor-library reference. Under the same budget, evaluated LLMs reach higher median instruction-qualified speedups with Triton than with CUDA-PTX, even without Triton's autotuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.