Reprise: Identity-Faithful Human Video Generation from Image and Video References
Abstract
Subject-driven video generation has made substantial progress in synthesizing videos of a referenced person in novel scenes, yet existing methods typically condition on a single reference image and measure identity by facial similarity between the generated face and that same reference. A reference captures the person under one specific viewpoint, styling, lighting, and expression, and is thus a sample of an identity rather than the identity itself, so same-reference protocols reward copying such transient attributes rather than preserving the person. The issue becomes more severe in multi-person scenarios, where no token-level cue binds each subject to its own references, leading to subject confusion and attribute swapping. We propose Reprise, a multi-reference identity-preserving video generation framework in which each character is specified by multiple image and video references from different contexts, and identity is defined as what remains invariant across these contexts. We construct Reprise-10K, a corpus of 10K identities derived from film and television, in which each identity is paired with multiple reference images and video clips spanning distinct viewpoints, lighting, expressions, and styling. We further introduce ID-RoPE, which extends rotary position embedding with an additional identity coordinate carried by every token, grouping attention by character so that all references of a character are attended jointly rather than any single one being favored, and avoiding the copying shortcut present in existing reference-conditioned designs. A cross-context identity reward post-trains the model by scoring generated faces against held-out references of each identity, steering the model toward generating the person itself rather than copying the conditioning references. Extensive experiments demonstrate that Reprise improves identity preservation, multi-person subject binding, and naturalness, while reducing identity confusion and excessive copying of the conditioning references.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.