acceptodds
Under review as a conference paper at ICLR 2027

I2VShield: Focused Adversarial Protection Against Harmful Image-to-Video Generation

Abstract

Modern image-to-video (I2V) diffusion models turn any photograph into a high-fidelity video, making non-consensual deepfakes cheap to produce. A natural defense is adversarial protection: adding an imperceptible perturbation (under a small budget) to an image before publication so that generation from it fails. Yet existing methods inherit the text-to-image (T2I) recipe, which barely dents modern I2V backbones. We trace this failure to the recipe's two pillars, neither paying off under I2V. First, the 3D VAE is far harder to move than its 2D counterpart, so the same budget buys far less disruption here. Second, and more decisively, I2V's static conditioning leaves random-timestep averaging both unnecessary and self-diluting: the perturbation never enters the noisy latent, and we measure cross-timestep gradients to be near-orthogonal (cosine I2V vs T2I ). Probing where the condition acts, we locate its influence at the largest noise level, decaying monotonically thereafter, so averaging spends most of the budget where the condition barely responds. Together, these findings identify what decides protection strength: not which objective is attacked, but where along the timestep axis the budget is spent. Guided by this diagnosis, we propose I2VShield, which forgoes the VAE objective and commits the entire budget to the largest noise level, where the condition's influence peaks. Across 9 I2V models and 2 datasets, it lowers DINO similarity from to on average, against for the strongest, outperforming even at a smaller budget ( vs DINO). We further release HQFace-I2V, the first benchmark purpose-built for this setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.