GRAFT: Video Unlearning by Grafting Frames Temporally
Abstract
Text-to-video (T2V) diffusion models are trained on large-scale uncurated corpora and inherit the ability to render unsafe concepts such as explicit nudity, public figures, or specific objects. Existing unlearning methods for video adapt image-domain recipes: they steer the model's prediction away from a concept prompt, along a direction that is not tied to any video the model would produce. We observe that video offers a source of supervision that images lack: because frames in one generated video share subject, lighting, and camera, a concept-free version of a concept-bearing video can often be obtained by substituting frames rather than generating new ones. We propose GRAFT, which grafts concept-free latent frames over concept-bearing ones and fine-tunes toward the edited latent with the model's own objective. Because most concepts are never partially present, we further propose split-prompt generation, which conditions different temporal regions of one latent on a concept prompt and its safe counterpart, manufacturing partial-concept videos for any concept. We evaluate GRAFT on datasets spanning nudity, ImageNet objects, and celebrity identities, showing stronger erasure than prior work on CogVideoX and comparable erasure on HunyuanVideo, with aesthetic quality close to the base model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.