Behave Yourself! Personalizing Talking Heads Beyond Appearance
Abstract
Recent video generation models enable appearance-based personalization by conditioning generation on reference images of a specific person. Yet a person's identity is not fully captured by appearance alone. It also includes characteristic behavior, such as distinctive head movements one makes while speaking or listening, which is not directly exposed through static reference images. In this work, we introduce a spatio-temporal identity encoder that aggregates a variable number of reference videos into a fixed number of identity tokens, capturing both visual details revealed across frames as well as the subject's characteristic behavior. These identity tokens condition a frozen audio-driven video diffusion transformer through dedicated cross-attention layers, enabling generation with new speech while preserving the subject's appearance and characteristic behavior. A three-stage training procedure first adapts the video encoder to talking-head videos through self-supervised training. It then establishes identity conditioning using the target clip as reference, before switching to different reference clips of the same subject. Through qualitative and quantitative experiments, we demonstrate that our method yields more authentic, person-specific videos compared to an appearance-only baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.