LLARA: Low-Rank Linear Attention Residual Adaptation for Large Language Models
Abstract
Modern large language models (LLMs) rely critically on residual connections with PreNorm, yet their uniform aggregation of layer outputs with fixed unit weights causes hidden states to expand uncontrollably with depth, gradually diluting the contribution of each layer. Recent research introduces Attention Residuals (AttnRes) to enable each layer to selectively aggregate earlier representations on a per-token basis through learned weights, by replacing the fixed accumulation with softmax attention over preceding layer outputs, yielding consistent empirical gains across model scales. However, pre-training AttnRes-based LLMs demands prohibitively high computational resources and expenses, preventing a wide range of open-source models from reaping the benefits of this novel architecture. To bridge this gap, we propose **L**ow-Rank **L**inear **A**ttention **R**esidual **A**daptation, or **LLARA**, which freezes the pre-trained model weights and introduces a learnable low-rank linear attention residual mechanism to fine-tune depth-wise aggregation of per-block residual contributions, allowing pre-trained LLMs to robustly benefit from token-wise attention residuals during downstream task adaptation. Jointly with LoRA for parameter-efficient fine-tuning, LLARA maintains a compact linear attention state along the decoder depth through low-rank key–value trajectories, reads it out at the final block via a learnable query, and adds the resulting residual back to the hidden representations, thereby fine-tuning the model’s residual architecture. Both theoretical analysis and empirical experiments have demonstrated the effectiveness, efficiency, and robustness of LLARA. Our code is available at https://anonymous.4open.science/r/LLARA.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.