acceptodds
Under review as a conference paper at ICLR 2027

HiLA: Historical-Layer Attention for Low-Rank Adaptation

Abstract

Parameter-efficient fine-tuning (PEFT) has become an important paradigm for adapting large language models (LLMs) to downstream tasks. However, existing methods such as LoRA typically perform local low-rank updates conditioned only on the current-layer input, leaving task-relevant representations from preceding layers underutilized. Our analysis shows that shallow and intermediate layers already contain information that can influence final predictions, motivating the explicit use of cross-layer representations for low-rank adaptation. To this end, we propose HiLA (Historical Layer Attention), a PEFT method that dynamically integrates cross-layer representational changes into low-rank updates. HiLA constructs differences between the current and preceding-layer states and uses the current LoRA bottleneck representation as a query to retrieve relevant historical information through attention. We evaluate HiLA on mathematical reasoning, commonsense reasoning, and open-domain dialogue generation across four backbones with different architectures and scales. HiLA achieves the highest average commonsense accuracy on all four backbones and the best performance on five of six dialogue-generation metrics. Notably, on challenging mathematical reasoning tasks, HiLA consistently outperforms all low-rank baselines and even surpasses full fine-tuning on Llama-3.2-3B with only 1.03% trainable parameters. These results indicate that attention-based cross-layer representations effectively alleviate the limitations of existing PEFT methods. Our code is available at https://anonymous.4open.science/r/HiLA-8209/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.