acceptodds
Under review as a conference paper at ICLR 2027

Seeing More, Refusing Less: Two Mechanisms of Safety Erosion in Video Large Language Models

Abstract

Expanding visual context in Large Vision-Language Models (LVLMs) exposes a capability–safety tension: more frames can improve video understanding while increasing unsafe responses to harmful video–text requests. The associated representational pattern depends on how harm is expressed. For explicit harmful queries, visual context narrows late-layer separation between harmful and neutral text representations, although risk remains decodable. By contrast, when benignly worded queries become harmful through video grounding, text-only risk cues are weaker, while failure-related state shifts align with an activation direction associated with unsafe responses. On selected inputs, amplifying the text-risk component or steering against this unsafe-response direction can lower judged-unsafe output rates. These findings motivate Risk-to-Refusal (R2R), a training-free defense that separates risk detection from refusal actuation. R2R routes requests using separate textual and visual risk signals, combining text-risk component amplification with refusal steering for explicit harm and applying refusal steering alone for grounded harm, with strength bounded as frame count grows. Across five LVLMs, R2R lowers attack success rates to 1.5–6.9% on explicit and 2.3–8.7% on implicit harmful requests.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.