acceptodds
Under review as a conference paper at ICLR 2027

Attention Output Projection Layers Are Strong Targets for Adaptation

Abstract

Effective adaptation need not update every component or add parameters. Tuning the right subset of existing parameters can be both more effective and more efficient. But where should a pretrained transformer be adapted? We conduct systematic comparisons of transformer components and identify the attention output projection as the strongest individual adaptation target. We attribute this to its role: attention output projections determine how attention features enter the residual stream, so when pretrained features are rich, recombining them can suffice. Our method, Attention Projection Layer Adaptation (APLA), tunes the output projection in each transformer block, or a fixed random subset of its columns, together with the task head. It adds zero parameters, requires no architectural changes, and adds no inference computation. On 21 classification datasets, we achieve state-of-the-art mean accuracy of 91.5% with DINOv2 ViT-B, outperforming 20 baselines (the strongest reaching 90.9%) including full fine-tuning (89.7%). Evaluation on 49 vision tasks in total supports these findings under low-data regimes, distribution shift, segmentation, and detection. With ViT-g, APLA reduces training memory by 53% and training time by 43% relative to full fine-tuning. Our finding extends beyond vision: output projections are the best attention component to adapt on six of eight GLUE language tasks. Lastly, we also find that parameter-adding methods such as LoRA and AdaptFormer benefit from placement at the output projection over their default location.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.