Exploiting Video Temporal Capability as an Attack Pathway in Multimodal Large Language Models
Abstract
Multimodal Large Language Models (MLLMs) are rapidly gaining advanced video understanding capabilities, including temporal localization, cross-segment reasoning, and long-context integration. However, existing video jailbreak attacks largely treat video as a rich carrier of harmful content, leaving these emergent capabilities underexplored as an attack pathway. In this work, we introduce **Video Capability-Safety Trap (VCST)**, whose key insight is to make harmful content recovery a prerequisite for completing a legitimate video understanding task, inducing an intrinsic conflict between capability-driven task completion and safety-driven refusal. VCST temporally disperses harmful information across video segments and requires the model to localize and integrate it, reducing frame-level exposure while co-opting the model's video reasoning capabilities to advance the attack objective. Extensive experiments on a diverse set of MLLMs show that VCST consistently outperforms prior video-based attacks across both non-thinking and thinking settings, and exhibits significantly stronger robustness to existing defenses. To counter VCST, we further develop a video-oriented defense that encourages deliberation on harmful intent detection. Additional analyses validate the reliability of our evaluation and examine the effect of model scale. Our findings highlight an emerging capability-safety tension in MLLMs and call for safety alignment that explicitly accounts for evolving video capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.