When More Capacity Is Not Enough: Rethinking Potential-Based Attention
Abstract
Potential-based attention models organize computation through a scalar function, providing mathematical structure for analyzing model behavior. A natural goal is to use this structure to reproduce an existing attention model, preserving its outputs on the same inputs while making its behavior easier to analyze. Contrary to the intuition that scaling model capacity can overcome such limitations, we find that an unavoidable output error can persist even with arbitrarily complex potentials. Generating multiple outputs through the gradient of a single convex function requires them to obey shared rules of variation, which the original attention response may not satisfy. For a class of fixed attention maps, we first prove a capacity-independent error lower bound that quantifies this structural restriction. We then give a criterion computable directly from the original model that determines the minimum number of independently corrected output directions needed to recover its outputs exactly. We also find that fewer corrections can sometimes approximate the original outputs arbitrarily well, but only as the relative scaling of output directions becomes increasingly uneven. This reveals the numerical cost of reducing the required structural change. Finally, numerical experiments with rigorous error guarantees establish that independent corrections can cross an accu- racy limit that increasing potential capacity alone cannot overcome. These results provide a theoretical basis for deciding when the architecture must change instead of continuing to scale the model. We believe these insights can guide attention designs that combine mathematical structure with expressive power.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.