ResiBand: Recovering Evidence Contrast under Local Common-Mode Dominance
Abstract
Many language-model tasks require identifying and combining a few pieces of relevant evidence in distractor-rich contexts. Models may use this evidence yet fail to distinguish it reliably from competing information. We investigate this difficulty in the residual stream. Controlled evidence-relocation experiments show that support–distractor differences remain present but are small relative to a dominant component shared by nearby residual-stream states. This imbalance is associated with weaker evidence discrimination in subsequent attention layers. We introduce ResiBand, a causal residual operator that subtracts a short-horizon exponential average, smooths the remaining variation over a longer horizon, and maps the filtered signal into a learned low-rank residual correction. The fixed filter exposes variation relative to the local background, while the learned mapping adapts the correction to the task with pretrained weights frozen. Under a shared configuration across backbones, ResiBand improves average F1 across three multi-hop QA benchmarks over LoRA by 19.30, 10.95, and 1.93 percentage points on Qwen3-1.7B, Qwen3-8B, and Llama-3.1-8B, respectively, while requiring fewer trainable parameters. Mechanistic analyses and residual interventions indicate that the learned corrections improve discrimination between supporting evidence and distractors. These results support temporal filtering and learned residual correction for better use of available evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.