Rethinking LoRA Placement at the Vocabulary Boundary with A-LoRA
Abstract
Low-Rank Adaptation (LoRA) is typically applied to internal Transformer layers, while vocabulary-boundary layers such as input embeddings and output LM heads have long lacked systematic evaluation. We revisit this placement question and propose A-LoRA, a hidden-space parameterization that does not learn free residuals directly on vocabulary matrices, but instead induces vocabulary-matrix updates via a shared low-rank affine transform, so parameter count depends mainly on hidden dimension rather than vocabulary size . Analyzing 30 Base→Instruct model pairs, we find that output LM heads and tied shared matrices better match this affine structure, whereas untied input embeddings should not be treated symmetrically under the same assumption. Downstream LoRA post-training further shows that optimal vocabulary-boundary placement is conditional: A-LoRA as a standalone small-budget adapter is stronger on the input side; when combined with standard hidden LoRA, the `lm_head` side becomes a more stable low-cost supplement. Our results indicate that vocabulary-boundary layers should not be excluded from LoRA design space by default, but their topology should adapt to existing adaptation capacity rather than being added symmetrically to input and output sides unconditionally.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.