acceptodds
Under review as a conference paper at ICLR 2027

OmniCustom: Sync Audio-Video Customization Using Contrastive Flow Matching

Abstract

Given a reference image and a reference audio , sync audio-video customization task aims to generate videos that maintain the identity of the reference image while imitating the timbre of the reference audio, with spoken content freely specifiable through user-provided textual prompts. While existing methods have achieved promising progress, a system that provides high-quality preservation of given references while yielding user-friendly inference latency remains a challenge. To deal with this, we propose OmniCustom, a powerful DiT-based sync audio-video customization framework that can synthesize a video following reference image identity, audio timbre, and text prompts all at once in a zero-shot manner. Our framework is built on three designs. First, identity and audio timbre control are achieved by modality-specific LoRA modules that operate through self-attention layers within the joint audio-video generation backbone. Second, we introduce a contrastive learning objective alongside the standard flow matching objective. It uses predicted flows conditioned on reference inputs as positive samples and those without reference conditions as negative samples, enhancing the model’s ability to preserve identity and timbre. Third, we train OmniCustom on our constructed large-scale audio-visual human dataset OmniCustom-1M. Extensive experiments demonstrate that OmniCustom achieves state-of-the-art customization performance while providing user-friendly inference runtime.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.