AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification
Abstract
High-performance ML systems increasingly rely on GPU kernels whose editable source is unavailable, generated, or too distant from final machine code to expose remaining optimizations. Existing LLM kernel optimizers and autotuners mainly operate on CUDA, Triton, HIP, or tensor-program source and validate against reference implementations. We study a stricter setting: optimizing an already compiled AMDGPU code object, using as the behavioral reference under a captured launch contract. We present AsmEvo, an agentic assembly-level optimizer for AMD GPU kernels. Given an AMDGPU code object , AsmEvo reconstructs a reassemblable representation, proposes profile-guided low-level edits with a long-horizon agent, rebuilds an ABI-preserving object with regenerated metadata or a conservative in-place patch, and accepts candidates only after capture-scoped equivalence and paired validation against in a generated persistent evaluator. On MI308X, AsmEvo improves all 30 selected KernelBench kernels, reaching geometric-mean and maximum speedup. On MI300X production workloads, it improves all evaluated AITER binaries and vLLM/SGLang Triton assembly kernels, reaching / and / geometric-mean/maximum speedups, respectively, while preserving functional equivalence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.