From Uniform Capacity to Functional Roles: Efficient Visual LoRA for Source-Free Cross-Domain Few-Shot Learning
Abstract
Source-Free Cross-Domain Few-Shot Learning (SF-CDFSL) aims to rapidly adapt pre-trained vision-language models to a target domain using only a few samples, without access to source domain data. However, existing vision LoRA methods typically employ a uniform Q/K/V adaptation structure across different Transformer layers, overlooking the distinct functional roles played by layers at different depths. This paper reveals an intriguing phenomenon regarding vision LoRA in SF-CDFSL: despite being assigned identical adaptation capacities, the layers contribute highly unevenly to the final decision. Specifically, shallow layers have a minimal impact, middle-to-deep layers rely primarily on the V-path, and final layers exhibit significant QK sensitivity. We investigate and explain this phenomenon; through counterfactual interventions on identical hidden states, we attribute this behavior to a stable functional division of labor emerging among Transformer depths and QK/V projection paths during target domain adaptation. Based on this insight, we propose Role-Aligned LoRA. This method aligns adaptation capacity with functional roles by freezing low-impact layers, retaining V-LoRA in V-sensitive regions, QKV-LoRA in transitional layers, and QK-LoRA in final layers. Experiments across four SF-CDFSL datasets validate our explanation and the proposed method. Results demonstrate that Role-Aligned LoRA maintains overall accuracy on CLIP while utilizing only 27.8% of the trainable parameters found in Uniform-QKV LoRA, and reduces support-set adaptation time by 34.3%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.