Beyond Multiplicity Correction: Sharp Retrieval Thresholds for Attention with Correlated Distractors
Abstract
Redundant context affects attention in two distinct ways: multiplicity adds normalization mass, while score variation within correlated groups can produce extreme individual scores. We separate these effects in a planted retrieval model with exponentially many Gaussian distractors organized into supplied clusters whose number and size both grow exponentially with dimension. The model is a two-level generalized random energy model, whose classical free energy yields exact retrieval thresholds. Gaussian interpolation further shows that increasing positive correlation at fixed token count lowers the distractor free energy, unlike adding duplicates. Our main result establishes a sharp limitation of count correction, which subtracts the log cluster size from pre-softmax scores: this correction removes multiplicity but retains within-cluster extremes. Even when the inverse temperature is chosen after observing all scores, count correction requires a signal at least , where is the exponential growth rate of the number of clusters. By contrast, averaging scores before exponentiation suppresses token-level fluctuations and lowers the optimal signal threshold to , where is the within-cluster correlation. This gap persists after adding exponentially many singleton distractors, so group size alone does not identify the target, and extends, under explicit cumulant regularity assumptions, to fixed finite hierarchies of independent-increment scores. Finite-sample, stability, and participation results complement the asymptotic analysis. A controlled retrieval experiment with frozen sentence embeddings illustrates both sides of the distinction: averaging helps when the relevant item is a singleton, but can discard useful evidence when a relevant group contains multiple matching items.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.