acceptodds
Under review as a conference paper at ICLR 2027

Omni-Customizer: End-to-End Multimodal Video Generation with Joint Identity and Audio Customization

Abstract

Despite recent advances in joint audio-video generation, precise appearance-and-voice customization remains challenging. We present Omni-Customizer, an end-to-end framework for human-centric joint audio-video customization. To establish cross-modal identity associations, we introduce Omni-Context Fusion (OCF) to integrate text embeddings with visual and audio references, together with Semantic-Anchored Multimodal RoPE (SA-MRoPE) to anchor paired references to their corresponding subjects. To reduce face duplication and timbre leakage under incomplete reference coverage, we further propose Subject-Aware Reference Routing (SARR), which controls where and how broadly each reference can directly influence generated content. We adopt a three-stage progressive training curriculum to achieve stable optimization and robust customization. We also develop a human-centric data curation pipeline to construct high-quality multimodal training data. Experiments on our proposed OC-Bench demonstrate strong video and audio quality, robust appearance-and-voice preservation, and faithful identity binding in complex multi-subject and incomplete-reference settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.