Attention's Gravitational Field: A Power-Law and Linguistic Interpretation of Positional Correlation
Abstract
This paper investigates the underlying principles of positional relationships and encodings in Transformer-based language models. We introduce the Attention Gravitational Field (AGF), which decouples positional encodings from semantic embeddings and models their interaction through a power-law distribution. Drawing on linguistic derivations, we demonstrate that AGF is intrinsically consistent with Newton's law of universal gravitation and his shell theorem, while also aligning with observed learning dynamics and stability curves. Empirically, AGF achieves consistent accuracy improvements over existing relative positional encoding methods on standard benchmarks. Our analysis provides a deeper perspective on the interpretability of attention mechanisms, revealing structural regularities in positional information that bridge linguistics, physics, and model design. This work constitutes a meaningful step toward understanding the fundamental laws underlying Transformer architectures.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.