acceptodds
Under review as a conference paper at ICLR 2027

STAR: Multi-Level Robustness Injection for Vision Transformer Downstream Adaptation

Abstract

Pretrained Vision Transformers (ViTs) are increasingly adapted downstream by parameter-efficient fine-tuning (PEFT) under a small learnable-parameter budget . Adapted models face transfer attacks from similar ViTs, and even white-box attacks if the learnable module leaks, so robust adaptation under such a is a practical need. Across PEFT methods and full fine-tuning, we find that standard adaptation is not robust and that adversarial training (AT) cannot balance clean accuracy and robustness once shrinks. We propose Smoothness and Token Alignment Regularization (STAR), a robust adaptation algorithm that spreads the robustness signal over multiple relaxed constraints that a small learnable subspace can satisfy: label-free class-token alignment, boundary smoothness, and down-weighted cross-entropies, at the representation, distribution, and decision levels, respectively. It applies unchanged to every adaptation method. Results from multiple datasets show that the advantage of STAR over AT is more pronounced under PEFT than under full fine-tuning. Analyses also show that robustness depends more on how learnable parameters reach the Encoder than on their number.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.