acceptodds
Under review as a conference paper at ICLR 2027

Associative-Retrieval Regime Entry Varies with Head Partition Under a Fixed Training Protocol in Transformers

Abstract

Partitioning a fixed attention space across heads can be associated with different associative retrieval outcomes under a shared training protocol. We train six-layer causal transformers with for 20,000 AdamW steps on key–value retrieval tasks using frozen sentence representations from WikiText-103 and AG News. On WikiText, we hold total attention width fixed at while varying the partition from 32 heads of dimension 16 to 4 heads of dimension 128. The fitted 50%-success boundary shifts from to associations, with the four fitted boundary estimates ordered across the tested partitions. A raw load selected on WikiText transfers poorly to AG News, motivating a separate AG News load map. We denote an architecture with heads of dimension by H-D. At the frozen AG News operating point of , H8-D64 enters the prespecified regime in of runs, compared with for H8-D32. An independent fixed-width comparison then matches total attention width, feed-forward width, parameter count, parameter shapes, initialization, data, and training budget. H8-D64 succeeds in of 120 runs, versus for H16-D32, yielding a paired risk difference of (95% CI ). These results show consistent differences in entry into the associative-retrieval regime across the tested head partitions under the shared training protocol.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.