acceptodds
Under review as a conference paper at ICLR 2027

AI as a Compiler: Compiling triton kernels without the triton compiler

Abstract

Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX, bypassing Triton’s intermediate representations and optimization passes while preserving the kernel ABI, and refines its output using randomized differential testing, benchmarking, and profiler feedback. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83×-3.34× the performance of autotuned Triton. The largest gains come from transformations that Triton’s lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34× on BitDelta), assigning each thread a complete softmax row in tensor memory (1.36× on FlashAttention), and reusing overlapping convolution windows (up to 2.23×). Getting correctness guarantees is now the bottleneck. The state-of-the-art PTX verifier does not model the Hopper and Blackwell features these kernels rely on, and restricting generation to its supported fragment erases the gains on compute-heavy kernels. With best- effort extensions, the verifier checks 16 of 22 kernels; atomics, input-dependent control flow, and bit-level floating-point manipulation remain out of reach, making verification the central open challenge for trusted AI compilation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.