EPO: Concept Erasing in Text-to-Video Models via Energy-based Erasing Preference Optimization
Abstract
Modern Text-to-Video (T2V) models generate increasingly realistic and coherent videos, but they also risk producing undesired content such as nudity, personal identities, or unsafe objects. Existing concept erasure methods, largely developed for Text-to-Image models, are difficult to transfer directly to T2V generation because undesired concepts can reappear through diverse appearances, motions, viewpoints, and temporal compositions. In this paper, we propose EPO, an energy-based preference optimization framework for T2V concept erasure. Instead of relying on local suppression of limited samples, EPO formulates erasure as energy-based preference shaping over the open-ended video generation space. We introduce an Energy-based Erasing Preference and instantiate it with an energy discrepancy objective that enlarges the energy gap between desired and undesired generations. To mitigate catastrophic forgetting caused by broad energy suppression, we further introduce a Manifold-guided Target Anchoring strategy, which redirects undesired prompts toward nearby safe semantic anchors while preserving spatiotemporal generation priors. Extensive experiments demonstrate that EPO effectively suppresses undesired concepts and better preserves video generation quality. Anonymous code is available at https://anonymous.4open.science/r/E-2PO-AnonymousRelease.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.