acceptodds
Under review as a conference paper at ICLR 2027

SABER: Semantic Abstraction Boundaries for Efficient Refinement of Sparse GPU Kernels

Abstract

GPU coding agents increasingly optimize kernels through search and execution feedback, but the programming interface presented to the agent remains underexplored. We present SABER, an agent-facing interface for sparse GPU kernels that exposes sparse strategy and coarse hardware intent while delegating concrete GPU realization to a compiler backend. Across 26 workloads spanning six sparse operator groups, SABER records no hard failures versus 11.1% for Direct CUDA, while 92.3% versus 47.4% of trajectories are within 10% of their final search best by proposal 2. Because the reachable implementation spaces differ, we treat this as a full-system comparison. We therefore construct a matched study on ten CSR SpMM/SDDMM workloads in which SABER and a lower-level schedule expose the same 32 executable backend implementations per workload. Across GPT-5.6 Sol, GPT-6 Astra, and MiniMax-M3, SABER reaches shared-reference thresholds earlier and selects implementations closer to the shared finite reference. Configured-backend and private-realization experiments further show that final kernel performance depends on backend quality. These results support treating the agent–compiler programming interface as a first-class systems design choice.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.