Reasoning Deeper through Controlled Sharpness
Abstract
Self-attention in Transformers uses softmax to normalize query-key products into a distribution of attention scores. A sharp distribution concentrates attention on a small number of tokens, whereas a flat one spreads it more evenly across tokens. Softmax, however, has no parameter that directly controls this sharpness. Instead, sharpness depends indirectly on the magnitudes of the query-key products, emerging through training and varying markedly across random seeds. We identify an empirical association between attention sharpness and reasoning depth under softmax. Motivated by this association, we introduce powered L1 attention (L1p), a normalizing function whose sharpness is controlled by a learnable per-layer exponent and whose attention scores are invariant to the overall magnitude of the query-key products. We further introduce a signed form of L1p that permits negative attention scores, which our ablations show is critical for reasoning depth. On a synthetic multi-hop reasoning task with an exactly controllable number of hops, L1p reaches greater reasoning depth than softmax and five established alternative normalizing functions within the same training budget. At hops, L1p solves the task in of runs with different random seeds, whereas every alternative succeeds in at most of . At the same time, across random seeds, L1p matches the language modeling perplexity of softmax within a pre-specified margin. These results show that directly controlling attention sharpness through an exponent, together with signed attention scores, can substantially increase the reasoning depth a Transformer achieves within a single forward pass.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.