Understanding and Improving Adversarial Answer Recovery via Activation Steering in Vision–Language Models
Abstract
Activation steering provides a controlled way to modify internal representations without updating backbone parameters. We study its use for *behavioral recovery*: restoring pre-attack responses after adversarial image perturbations while preserving responses to pre-attack inputs. Controlled analyses first examine where and how to intervene. We find that attack status is decodable at the answer-prediction state, yet detector and recovery directions are weakly aligned, and input-specific answer-margin directions recover substantially more responses under matched interventions. These findings motivate AM-GARD, which separates intervention selection from answer correction using a learned intervention gate and input-specific answer-margin updates. Across multiple LVLMs, tasks, and answer formats, AM-GARD improves recovery while maintaining high pre-attack response preservation. Further analyses show that broader selection trades preservation for recovery opportunity, while successful correction depends on identifying a useful answer boundary rather than merely changing the output. Together, these results show that behavioral recovery depends on separating *whether* to intervene from *how* to correct the answer, providing both a diagnostic account of recovery behavior and an effective inference-time intervention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.