Vorch-IR: Long-Form Multimodal Identity Replacement Video Generation via a Unified Framework
Abstract
Identity replacement aims to replace subject identities in videos while preserving motion, expressions, and temporal consistency. Existing methods primarily focus on single-person scenarios and rely on task-specific structural guidance, such as masks, pose maps, or motion control signals, making them difficult to integrate into unified image-text-video editing frameworks. Moreover, the lack of large-scale paired training data limits progress on multi-person replacement. We propose a unified identity replacement framework that supports both single-person and dual-person replacement, with optional background replacement, within a single model. Built upon the LTX2 backbone, our method jointly conditions on textual instructions, reference identity images, optional background images, and driving videos. Notably, our approach eliminates the need for pose or spatial alignment in the reference inputs; multiple subject identities and background references can be fully decoupled and directly fed into the model. To support training, we develop an automatic data construction pipeline that synthesizes high-quality paired data for all editing settings. We further extend the model to long-form video generation and introduce Long-Horizon Error Control, enabling minute-scale video generation. Experiments demonstrate strong identity preservation, motion consistency, and temporal coherence across diverse replacement scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.