acceptodds
Under review as a conference paper at ICLR 2027

Counting with Attention and LayerNorm: Norm Scaling and Representation Geometry

Abstract

Counting is a simple task on which transformers can nevertheless fail at long sequence lengths. We study the cost of counting in a single attention layer, up to a maximum length , focusing on how softmax attention and layer normalization shape the representation from which the count is recovered. We mainly focus on regression read-out but also consider classification read-out. For regression, we show unnormalized linear attention counts exactly at parameter norms independent of . For softmax attention, regression cannot count exactly with finite parameters; it recovers the count after rounding, but only when the vector parameter norms grow with . We measure this cost by , the smallest bound that can be placed on every trainable vector to ensure exact counting after rounding, and derive tight rates. At fixed model dimension, grows as without normalization, as under bare post-normalization, and as with learned post-normalization. The last rate is unexpected: post-normalization discards the magnitude that carries the count and might be expected to make counting harder, yet with a learnable gain it counts at a smaller norm than softmax without normalization. Bare pre-normalization is different: it bounds the attention scores, so beyond a finite length no parameter norm suffices. A classification read-out keeps every norm bounded under every placement, but a dense head stores one output vector per count and costs more in storage. Experiments on the post-normalized model reproduce this split between the two read-outs. Together, our theoretical and empirical findings characterize how self-attention, layer-normalization placement, and read-out choice shape the representation geometry and parameter cost of counting in the minimal single-head architecture studied here.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.