Attention Regulates Resolution, Not Sharpness: A Diffraction Limit for the Softmax Scale
Abstract
While the canonical softmax scale preserves variance at initialization, it fails to capture the operating regime of trained attention. With normalized queries and keys, logits factor exactly into a cosine similarity and an explicit scale , where the canonical choice corresponds to setting . Because the learned geometry can spread its cosines to offset an imposed scale, the scale and the geometry are redundant; we sweep on a logarithmic grid to examine how autoregressive training resolves this redundancy. Training absorbs much of the imposed scale into the learned geometry, yielding a rank correlation of across a 32-fold sweep, and conserves a logit margin rather than a scale. Indeed, eight public models with learned scales ranging from 11 to 73 maintain nearly the same margin. At the optimum, this margin sits at or just past the diffraction limit, defined as the threshold at which a single key claims half the attention weight against equal competitors, a property we formalize and verify in Lean 4. Thus, while the scale dictates sharpness, the margin normalized by the diffraction limit dictates resolution, and next-token prediction regulates resolution rather than sharpness. Empirically, the loss-optimal scale sits at from 25M to 393M block parameters, across context lengths from 512 to 2048 tokens and vocabularies from 6 to 32,768 tokens, on text and RNA, with one slight exception. At this fixed scale, 64-dimensional heads nearly match 128-dimensional heads, and learned global, per-layer, or per-head scales fail to outperform the constant value. Degradation under the canonical grows quadratically with its log-distance from this optimum, a distance set by the head dimension alone, and steepens as the context lengthens. These findings show that comparisons across head sizes under the canonical scale conflate a systematic scale artifact with architectural capacity, and that freezing decouples attention resolution from head size.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.