acceptodds
Under review as a conference paper at ICLR 2027

A model-agnostic scaling rule behind the critical inverse temperature in softmax self-attention

Abstract

Length-dependent logit rescaling is widely used to stabilize long-context self-attention, but existing analyses and methods suggest different inverse-temperature laws for the context length , ranging from to and . We provide a general theory showing that the critical scale is determined by the gap-counting function of each attention row, without specifying the details of the model structure. Counting how many competitors lie within each gap from the maximum, we define an upper-tail accumulation scale and prove that it gives the critical inverse-temperature scale for softmax concentration: below this scale, the top competitors remain unseparated, whereas above it, the attention entropy collapses. This framework provides a model-agnostic, higher-level explanation of how scaling laws vary with , and yields a direct diagnostic for attention-score families, from idealized theoretical models to practical transformers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.