Can Visual Prompts Inherit Full Fine-Tuning?
Abstract
Pretrained vision transformers have demonstrated strong transfer ability and are increasingly shared across many downstream tasks. This trend has driven growing interest in parameter-efficient adaptation, which aims to adapt one frozen backbone to each task with only a few task-specific parameters. Visual prompt tuning (VPT) is a representative approach that learns a small set of prompt tokens and a classifier for each task. However, existing VPT methods learn prompts only on the frozen pretrained backbone and therefore cannot use the task-specific adaptation that full fine-tuning (FT) provides. These limitations leave compact prompts well below the accuracy their architecture can reach. Therefore, this paper proposes gradual withdrawal. Specifically, gradual withdrawal first fine-tunes a copy of the backbone on the target task, and then trains visual prompts while smoothly returning the backbone to its original pretrained weights. The deployed model is exactly that of direct VPT: the original backbone plus a prompt and a classifier per task. Crucially, this ordered return lets the prompt track a sequence of nearby solutions, as our local tracking analysis explains. On supervised ViT-B/16, gradual withdrawal raises mean accuracy from 69.4% to 76.2% across 19 VTAB-1k tasks and from 89.1% to 92.3% across five fine-grained classification tasks, although the FT model itself scores only 65.6% and 88.5%. Extensive experiments show that gradual withdrawal outperforms matched distillation, abrupt-return, and shuffled-path controls, with gains that hold across eight backbones spanning 86M to 1.1B parameters. More broadly, these results establish the training path as a new axis of parameter-efficient adaptation: a compact model can learn from far more than it deploys, so better training alone can make it stronger.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.