Seeing More, Refusing Less: Two Mechanisms of Safety Erosion in Video Large Language Models
Abstract
Expanding visual context in Large Vision-Language Models (LVLMs) exposes a capability–safety tension: more frames can improve video understanding while increasing unsafe responses to harmful video–text requests. The associated representational pattern depends on how harm is expressed. For explicit harmful queries, visual context narrows late-layer separation between harmful and neutral text representations, although risk remains decodable. By contrast, when benignly worded queries become harmful through video grounding, text-only risk cues are weaker, while failure-related state shifts align with an activation direction associated with unsafe responses. On selected inputs, amplifying the text-risk component or steering against this unsafe-response direction can lower judged-unsafe output rates. These findings motivate Risk-to-Refusal (R2R), a training-free defense that separates risk detection from refusal actuation. R2R routes requests using separate textual and visual risk signals, combining text-risk component amplification with refusal steering for explicit harm and applying refusal steering alone for grounded harm, with strength bounded as frame count grows. Across five LVLMs, R2R lowers attack success rates to 1.5–6.9% on explicit and 2.3–8.7% on implicit harmful requests.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.