FlowFaceSwap: Training-free Video Face Swapping via Structure-Preserved Identity Flow Editing
Abstract
High-fidelity video face swapping requires transferring a source identity seamlessly while strictly preserving the target's expressions, poses, and temporal consistency. Achieving satisfactory swapping performance typically necessitates resource-intensive, task-specific training, making these models expensive and inherently limiting their generalization to unseen scenarios. Recent training-free approaches have emerged as promising alternatives and achieved notable progress. However, they still struggle with temporal inconsistencies and identity leakage in face swapping. In this work, we reformulate the face swapping task as an identity-guided flow residual generation process, and propose a training-free framework FlowFaceSwap based upon a frozen rectified flow prior. Specifically, by optimizing compact identity tokens at test time, our method establishes a discriminative identity context. To prevent the overwriting of target dynamics, we introduce a structure-aware residual modulation mechanism driven by temporally tracked semantics and landmark-guided gates. This spatial control accurately localizes identity transfer while rigorously maintaining the target's subtle expressions and fine-grained motions. Extensive experiments on FaceForensics++ and CelebVHQ demonstrate that FlowFaceSwap generates temporally stable, highly realistic results. It significantly outperforms existing training-free baselines in both visual quality and identity retention, achieving an overall performance highly competitive with specialized, training-based models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.