A Narrow Residual Bottleneck from Evidence to Answer in Language Models
Abstract
When a language model answers a question about a passage, the answer hinges on a small span of the input. We ask where inside the model that information is forced to pass, and find a narrow residual bottleneck: a small set of evidence-token positions at a single residual-stream layer that carries the evidence-to-answer computation on its own. We locate it with two hidden-state interventions on a single candidate set. One copies the original evidence into a corrupted run and asks whether the correct answer returns (sufficiency); the other injects the corrupted state into a clean run and asks whether the answer breaks (necessity). Requiring the same set to pass both is strictly stronger than either direction alone. Against a size-matched permutation null, the existence of a bottleneck becomes falsifiable: a bottleneck must pass the two-sided test far more often than random sets of its own size, and it does. On the controlled benchmark, in our best-powered models the cut passes about 0.98 of the time against about 0.01 for the matched control, and is narrow (typically two to three positions, one to two on real text) throughout. It reappears across three normalization schemes on real single-hop, multi-hop, arithmetic, and long-context question answering, and at a fixed budget steers the answer far more reliably than editing a random set. Most tellingly, the bottleneck is not simply the top of an importance ranking: fixing the set size, the test, and the model, and changing only how the set is chosen, the same-size set of highest-importance attention heads fails the very test the located cut passes, a gap that holds across four model families.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.