acceptodds
Under review as a conference paper at ICLR 2027

Styling Past the Guard: Content-Preserving Style Attacks on VideoLLM Pipelines

Abstract

Visual style can alter the safety behavior of vision-language models without necessarily changing the content they recognize. In guarded Video Large Language Model (VideoLLM) pipelines, however, a successful attack must bypass an independent input content guard while preserving the harmful event that originally triggered detection. We present CTASO, a constrained temporal adversarial style optimization framework that searches a continuous space of video transformations under explicit content-preservation constraints. CTASO combines temporally coordinated editing, source-based content restoration, and feasibility-first optimization using input-guard feedback without querying downstream VideoLLMs. We evaluate CTASO across three content guards and four VideoLLMs, comparing it with five baselines and five adapted prior attacks. Evaluation is restricted to videos initially blocked by each guard, and event preservation is independently verified through human audit. CTASO consistently outperforms budget-matched random style search, achieving a human-audited Guard Evasion Rate of 0.61 versus 0.22 on LlavaGuard. Its effectiveness also extends to downstream harmful generation after excluding transformations that fail human audit. Comparisons with prior attacks reveal that effectiveness against unguarded models does not necessarily translate into success against a guarded pipeline. Moreover, apparent guard evasion can result from transformations that compromise the original harmful event. These findings establish content-preserving video style transfer as a distinct attack surface and highlight the importance of jointly evaluating guard evasion, event preservation, and downstream safety alignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.