acceptodds
Under review as a conference paper at ICLR 2027

Extrapolation Collapse Is a Readout Failure: Long-Context Models Retrieve What They Cannot Say

Abstract

Interpolation methods such as YaRN let a pretrained model be evaluated far beyond its training window, and they are judged by end-task score. By that measure they collapse: YaRN-extended Qwen3-8B falls from 0.51 at 64K to 0.23 at 128K on a controlled retrieval task. We show that the retrieval computation largely survives this collapse and that the failure lies in the readout. Letting 20 frozen retrieval heads vote over the parsed candidate set, without gold labels and without changing any weight, recovers the correct answer in 71 to 96% of the samples the model gets wrong, at every extension depth we measure, and never overturns a correct one. We call this gap the *readout bottleneck*, and it breaks behavior-only evaluation: at 128K one checkpoint scores 0.00 while locating the answer in 92% of its failures, and the behavioral ranking of two checkpoints inverts relative to their retrieval ability. Cutting the output path into four stages of decreasing dependence on the token distribution places the damage after retrieval rather than in it, and the damage itself is confined to the frequency bands the extension method rewrites, which we predict analytically before measuring. Its mechanism is de-anchoring to filler text rather than confusion among candidates. The bottleneck also bounds what interventions can do: steering the distribution with the same attention signal, including a published method run under our protocol, moves behavior by at most 0.06 where reading the trace moves it by 0.4 to 0.5. These results invite a reconsideration of what long-context benchmarks measure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.