ROUTING SHUFFLES DO NOT IDENTIFY ADAPTIVE BENEFIT: THEORY AND A CONTROLLED ACM-GCN AUDIT
Abstract
When a trained network loses accuracy after its routing gates are shuffled, it is natural to attribute the loss to the value of adaptive routing. We show why this inference requires care. Shuffling changes the assignment of gates while preserving their collection; replacing them by a constant changes both. For separable gate-to-output computation and additive node scores, the expected shuffled score equals an average of constant-gate scores. The shuffle drop therefore upper-bounds the frozen network’s advantage over the best constant gate. Once message passing couples the gates, either ordering is possible, as two explicit constructions show. We examine the distinction in ACM-GCN on five previously studied graphs, using 2,100 tuning fits and 750 final models. On Roman-empire, the joint shuffledrop is 11.59 accuracy points, the gap against a searched frozen constant is 4.27 points, and the advantage over a separately trained constant-routing model is 1.42 points. These comparisons concern different properties of the network. Our analysis explains when a shuffle test admits a bound and why a practical assessment of routing should also include searched constant gates and separately trained static controls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.