Beyond Static References: Audio-Visual Reference Conditioning for Human Video Generation
Abstract
Reference-conditioned human video generation has advanced rapidly, yet existing datasets still predominantly represent human subjects using one or a few static reference images. Such references provide only sparse observations of subject appearance while discarding temporal visual cues and voice characteristics naturally available in videos. Large-scale training data with richer audio-visual references for human video generation therefore remain limited. We introduce VidVoice-Ref, a large-scale dataset for human video generation with audio-visual subject references. Instead of reducing a subject to isolated frames, VidVoice-Ref preserves complete reference videos together with synchronized audio, capturing multi-view appearance, pose and expression variations, motion patterns, and voice timbre. We construct identity-consistent cross-video samples by identifying target subjects, retrieving temporally separated reference clips, and selecting visually diverse references across viewpoints, poses, expressions, motions, and scenes. Audio references are further refined through speaker diarization, speech validity filtering, intra-clip speaker consistency, voice quality filtering, and cross-video voice matching. VidVoice-Ref contains approximately 1M reference–target samples and over 1K hours of video, covering both single- and multi-subject scenarios. VidVoice-Ref provides a scalable data foundation for studying richer human subject representations beyond static image references, particularly for identity preservation, subject binding, and audio-visual reference-conditioned human video generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.