acceptodds
Under review as a conference paper at ICLR 2027

Faster Generalized Message Passing with Skew-Aware Parallelism and Work Partitioning

Abstract

Graph neural networks (GNNs) are built from sparse operations that gather and aggregate features along the edges of a graph, and on GPUs their speed is limited by memory access and parallelism rather than by arithmetic. IO-aware kernels remove most of the data movement, yet still leave much of the GPU idle on real graphs and cover only a few fixed layers. We show that the remaining bottleneck is how work is partitioned: degree distributions are heavy-tailed, so a kernel that gives every node one unit of parallel work waits on its few highest-degree nodes while the rest of the device idles. We, therefore, make GNN kernels degree-aware. First, nodes are bucketed by degree, extending a technique of prior IO-aware kernels, then reordered within each bucket and processed concurrently, so that long jobs start first and short ones fill the idle capacity. Second, the heaviest neighborhoods are cut into fixed-size edge slices that are merged exactly, and per-edge computations are partitioned over edges entirely. Third, data is loaded asynchronously wherever there is computation to overlap it. We bring these techniques to the two generalized message-passing primitives, gSpMM and gSDDMM, and provide them as a portable library of and kernels with autograd, from which arbitrary GNN convolutions can be composed in pure PyTorch. Against DGL, a widely used and efficient graph learning library, we achieve speedup for GATv2, speedup for gSDDMM, and for gSpMM in the median over thousands of configurations, reaching on and on . Our implementations are also faster than PyG and can process graphs which cause other libraries run out of memory. Against prior IO-aware kernels, they are faster in the median and on the most skewed graph.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.