APPLE: Recovering Attention-Attenuated Evidence with a Low-Rank Residual Path
Abstract
Softmax attention forms contextual representations by taking a weighted average of value vectors across tokens for each query, but this aggregation can attenuate token-specific distinctions that remain useful for prediction. We introduce APPLE (Attention-Preserving Path for Local Evidence), a lightweight low-rank branch that can be added to Transformer attention. Given contextual output , APPLE adds a low-rank correction from the local–context residual , where denotes query-aligned local values and equals in standard self-attention. This gives downstream layers access to fine-grained, token-specific information that can be attenuated in the attention output, while retaining the contextual representation. Under the official 300B-token Pythia training recipe, APPLE improves best-validation perplexity at every evaluated scale from 70M to 410M, with reductions of 6.04%, 9.69%, and 0.42%, respectively. The models also improve or tie on most of the 12 zero-shot benchmarks. Complementary controlled experiments show that the attention residual carries label-relevant signals attenuated by attention aggregation, while branch-off and residual-shuffling interventions in trained vision and language models degrade predictions. Across nine prediction settings spanning language, vision, multimodal retrieval, video, audio, and 3D synthesis, APPLE improves matched performance. These results support low-rank access to local–context residuals as a lightweight way to improve prediction while retaining contextual mixing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.