acceptodds
Under review as a conference paper at ICLR 2027

CleanVideo: Adaptive Subspace Intervention for Text-to-Video Concept Erasure

Abstract

Concept erasure in text-to-video diffusion models requires interventions that follow a concept as it emerges across spatial locations, video frames, and denoising steps. Fixed interventions can miss transient concept appearances or suppress unrelated content, compromising both erasure and temporal coherence. We introduce CleanVideo, a framework for selective concept erasure through adaptive interventions in low-dimensional hidden-state subspaces. A gating module jointly conditions on spatiotemporal visual features, denoising timesteps, and text semantics to control the location and strength of each intervention. We train the intervention modules with the pretrained diffusion backbone frozen, using target–surrogate prompt pairs to redirect unwanted concepts and preservation objectives to retain benign content. Experiments on CogVideoX-2B, CogVideoX-5B, and HunyuanVideo cover nudity, artistic style, and object erasure. Across these settings, CleanVideo improves the trade-off between concept suppression and generation quality relative to the evaluated baselines, while preserving visual fidelity and temporal coherence. Evaluations further measure video leakage whenever a target is detected in any sampled frame and assess concept recovery attacks with the intervention modules kept active. These results support adaptive representation intervention as an effective approach to selective video concept erasure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.