HyRA: Head-wise Reflective Adaptation via Input-Dependent Householder Transformations
Abstract
Low-rank adaptation (LoRA) has become a dominant approach to parameter-efficient fine-tuning, which is widely used in large foundation models. LoRA can be conceptualized as projecting the input space into a low-dimensional latent adaptation space, where the dimensionality is determined by the rank of LoRA. However, under multi-head attention, standard LoRA constrains all heads to share a single latent adaptation space, despite attention heads typically encoding functionally distinct roles in pretrained models. We identify this mismatch between the shared adaptation and head-wise specialization as a structural bottleneck that limits LoRA's expressiveness. In this paper, we propose Head-wise Reflective Adaptation (HyRA), which dynamically generates an input-dependent Householder transformation for each head. This Householder generator enables compact head-wise adaptation by modulating feature components along an input-dependent direction while preserving the orthogonal residual. Formally, the head-wise HyRA update can be expressed as , where and are low-rank matrices and is a head-wise, input-dependent Householder transformation that breaks the shared subspace bottleneck. Notably, HyRA does not increase the LoRA rank, but enables each head to learn its own input-dependent adaptation. Extensive experiments across both language and vision architectures demonstrate that HyRA consistently outperforms LoRA and its competitive variants across diverse downstream benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.