acceptodds
Under review as a conference paper at ICLR 2027

Vision Encoder Is a Hidden Culprit for Safety Degradation in Benign LVLM Fine-Tuning

Abstract

Benign fine-tuning can undermine the safety alignment of Large Vision-Language Models (LVLMs), a risk commonly associated with their language counterpart. We reveal that the vision encoder is also unsafe or less safe than the language backbone. By updating only the vision encoder (V-tuning), substantial safety degradation is observed across various LVLMs and downstream tasks. Such safety deterioration matches or even exceeds that of language-backbone or full-model tuning. To explain this vulnerability, we propose the vision-sensitivity hypothesis and analyze how parameter updates perturb representations along a fixed refusal direction. Our first-order analysis finds that the refusal-mediating structure in LVLMs is significantly more sensitive to the vision encoder updates. Based on this, we then demonstrate that V-tuning substantially reduces the separation between harmful and benign representations along the refusal direction more than the other tuning modes. Motivated by these findings, we introduce SAVR, Sensitivity-Aware Vision Regularization, to preserve the LVLM's refusal structure during vision encoder fine-tuning. SAVR successfully reduces attack success rates by up to 98% relative to unprotected V-tuning while preserving utility. It also retains most of the aligned model's harmful-benign separation. We believe that this work offers new insights for further studies to understand and address fine-tuning risks in LVLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.