acceptodds
Under review as a conference paper at ICLR 2027

The Shadow of a Few Keys: Local Low-Dimensional Structure in Long-Context Attention

Abstract

Efficient long-context attention rests on an implicit chain of assumptions: a low-dimensional query–key interaction should admit a small separable representation, that representation should give cheap and accurate attention, and accurate attention should preserve the model's answers. We test this chain link by link in two frozen language models, using local coordinates adapted to each block of queries. The first link fails: better local directions sharply reduce the error of an exact-exponential diagnostic, yet finite-state kernel summaries on the same coordinates recover only a small part of that gain. We call this discrepancy the geometry–realization gap and then locate it. A small set of high-mass keys, under one percent of the history, suffices to rebuild nearly all of the useful local geometry, and computing exactly on those same keys closes most of the gap, an effect that survives controls for projection error, attention sinks and the choice of support. We call this the few-key shadow. Two boundaries follow. A low-dimensional remainder often pays off under matched storage but never under matched arithmetic in our tests, and the support that reconstructs average geometry is far smaller than what the answer-determining queries need, so a frozen fresh-input evaluation still falls short of task preservation on three models. The results separate geometric compressibility, computational realization and task fidelity, and show that the first does not certify the other two.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.