acceptodds
Under review as a conference paper at ICLR 2027

Vorch-IR: Long-Form Multimodal Identity Replacement Video Generation via a Unified Framework

Abstract

Identity replacement aims to replace subject identities in videos while preserving motion, expressions, and temporal consistency. Existing methods primarily focus on single-person scenarios and rely on task-specific structural guidance, such as masks, pose maps, or motion control signals, making them difficult to integrate into unified image-text-video editing frameworks. Moreover, the lack of large-scale paired training data limits progress on multi-person replacement. We propose a unified identity replacement framework that supports both single-person and dual-person replacement, with optional background replacement, within a single model. Built upon the LTX2 backbone, our method jointly conditions on textual instructions, reference identity images, optional background images, and driving videos. Notably, our approach eliminates the need for pose or spatial alignment in the reference inputs; multiple subject identities and background references can be fully decoupled and directly fed into the model. To support training, we develop an automatic data construction pipeline that synthesizes high-quality paired data for all editing settings. We further extend the model to long-form video generation and introduce Long-Horizon Error Control, enabling minute-scale video generation. Experiments demonstrate strong identity preservation, motion consistency, and temporal coherence across diverse replacement scenarios.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.