Attribution Sensitivity Guided Parameter Allocation for Adversarially Robust Parameter-Efficient Fine-Tuning
Abstract
Pre-trained foundation models that undergo standard pre-training can be efficiently fine-tuned for downstream tasks using parameter-efficient fine-tuning (PEFT) methods. However, these models remain highly vulnerable to adversarial perturbations. Although existing works provide solutions for adversarially robust PEFT, they often distribute PEFT parameters uniformly across layers, and less focus has been given to the distribution of the limited capacity of PEFT methods across model layers tailored for adversarial robustness. In this work, we introduce a novel method that guides the allocation of LoRA ranks and parameters across transformer blocks, providing better adversarial robustness compared to uniform allocation. Specifically, we re-examine a representation-level diagnostic, attribution sharing, for robust PEFT, which reveals mixed alignment between the gradient of the adversarial loss and attribution sharing. We then provide a theoretical analysis of why this mixed alignment can impede adversarial training optimization. This motivates our method, Attribution Sensitivity, a one-shot method that assigns more LoRA capacity to blocks where global attribution sharing is less sensitive to relative weight changes. Experiments on CIFAR-10 and ImageNet-R with standard-pretrained ViT-B/16 and DeiT-Tiny backbones across three rank budgets achieve the highest average robust accuracy, supporting attribution sensitivity as a useful signal for robustness-oriented LoRA rank allocation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.