Receiver-Query Diversity in Graph Transformers: Frozen Interventions Misestimate Adaptive Capacity
Abstract
Full attention is often valued for its receptive field. Here I investigate a different resource: how many distinct receiver-specific summaries a model can construct from that shared context. Whereas global pooling broadcasts a single uniform representation to all nodes, full attention can tailor a report to each node. I formalize the spectrum between these extremes by parameterizing attention through discrete receiver-query modes. Theoretically, a random-access lower bound and matching construction establish the necessity of query diversity, while controlled graph experiments demonstrate that the demand diversity of receiver tasks, rather than context size alone, governs the required number of modes. Empirically, restricting query diversity within standard Graph Transformers significantly degrades dense node-level segmentation (COCO-SP, PascalVOC-SP), confirming that query diversity represents a primary capacity bottleneck. Crucially, I show that standard frozen ablations, commonly used to assess architectural components, fundamentally misestimate the capacity recoverable through end-to-end adaptation. Across multiple benchmarks, training with constrained queries recovers substantial performance over post-hoc frozen interventions, yet this recovery gap non-monotonically reverses at higher capacities. The findings establish that receiver-query diversity is a genuine capacity axis for Graph Transformers, while demonstrating that frozen sensitivity, adaptive capacity, and routing trainability are distinct properties that standard ablation practices routinely conflate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.