acceptodds
Under review as a conference paper at ICLR 2027

Multi-View Identity-Consistent Video Generation under Large Facial-Angle Variations without Cross-Paired Data

Abstract

Reference-to-video generation still struggles to preserve a known identity when the generated face turns far from a single reference view. Adding multiple views supplies missing identity cues, but in-paired training introduces a complementary failure: the model can copy one reference angle for several frames and then switch abruptly to another. We call this failure view-dependent copy-paste. We present Mv²ID, a multi-view conditioned framework that improves identity consistency while retaining natural facial motion using only in-paired supervision. The method concatenates video and reference latents, applies region-masking training to prevent direct copying and encourage complementary cue aggregation, and uses a reference-decoupled rotary positional encoding (RD-RoPE) so reference images are not interpreted as future video frames. We also build a 22K-video large-angle dataset and introduce Multi-view Reference Consistency (MvRC) for identity evaluation across viewpoints. On a held-out set of 30 identities and 150 prompts, Mv²ID obtains MvRC scores of 0.544/0.507 and a NaturalScore of 4.69, achieving the best identity consistency and text–video alignment among the compared methods while remaining competitive in naturalness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.