Distance is More Than a Bias: Geometry-Driven Value Mixing in Graphs
Abstract
Geometric deep learning and graph transformers have driven significant advancements in learning from structured data, particularly in molecular property prediction and spatial modeling. However, while recent approaches often incorporate spatial distances as attention biases or complex pair representations, this significantly increases computational overhead. In this paper, we propose a paradigm shift: rather than treating distance merely as a bias, we utilize the distance matrix as a direct value-mixing operator to achieve both high predictive performance and extreme computational efficiency. We introduce DisAttention, which replaces standard query-key attention with mixing weights derived from learnable transformations of the distance matrix. To further enhance expressive power, we propose DisMix, which integrates node feature affinity through symmetric normalization and distance gating instead of conventional row-wise softmax. By eliminating the need to update high-dimensional pair states, our geometry-driven approach drastically simplifies computation. Extensive experiments across molecular datasets (QM9, MolPCBA, LP-PDBbind) and broader geometric domains (MNIST superpixels, ModelNet40) demonstrate that our models achieve highly competitive accuracy while significantly reducing parameter count, peak memory usage, and training time compared to state-of-the-art graph transformers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.