acceptodds
Under review as a conference paper at ICLR 2027

Seeing Sparsity: Pattern-Guided Agents for Hardware-Adaptive SpMV Optimization

Abstract

Sparse matrix-vector multiplication (SpMV) performance depends on both matrix sparsity and GPU architecture, making it difficult to choose an effective kernel from structure alone. We present Sparrow, which combines a sparsity-aware decision model with feedback-guided kernel search. The model encodes local nonzero occupancy patches with a Transformer and global matrix statistics with an MLP to rank expert-defined strategies for work mapping and reduction. An agent starts from these strategies, tunes template parameters, and generates or revises kernel code using compilation, correctness checks, target-GPU measurements, and available profiler feedback in a search tree. Within its search budget, it returns the lowest-latency kernel that passes its correctness checks. We evaluate Sparrow on NVIDIA A100 and Hygon DCU GPUs against vendor libraries, a supervised strategy selector, and general-purpose code agents. With GLM, Sparrow achieves a speedup over cuSparse on NVIDIA A100 and a speedup over hipSparse on Hygon DCU. In ablations, throughput falls from 207.38 GFLOPS to 190.33 without sparsity awareness and 134.59 without the agent search. The results show that structure-based strategy selection and target-device feedback both contribute to SpMV kernel performance across the evaluated GPUs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.