Causal-Role Erasure for Text-to-Video Diffusion Models
Abstract
Video object removal is an important capability with broad applications in controllable generation and content editing. However, erasing an object does not always erase its physical causal effects: a droplet may vanish while its splash remains. Existing text-to-video concept-erasure methods target named concepts rather than their causal roles, leaving this conditional removal without an explicit formulation or training strategy. We formulate causal-role erasure as suppressing a source and its effects when it causes an event while preserving the same object's noncausal uses, and introduce Source-slot Randomized Counterfactual Distillation (SRCD) to address it. Specifically, we construct independently generated and screened source-free target videos that retain the receiver and scene without the source or its effects, then pair each target with event prompts that vary only the source identity. Across seven physical mechanisms, SRCD obtains a higher average graded erasure score than an otherwise matched control, but the corrected difference is not statistically significant; paired-role tests show no clear additional benefit, and human review does not confirm the automated comparison.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.