acceptodds
Under review as a conference paper at ICLR 2027

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

Abstract

Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synchronization. Reinforcement Learning (RL) offers a promising paradigm, but its extension to and joint audio-video generation remains unexplored. Notably, our in-depth analysis first reveals that the primary obstacles to applying RL in this stem from: multi-objective advantages inconsistency, where the advantages of multimodal outputs are not always consistent within a group; multi-modal gradients imbalance, where video-branch gradients leak into shallow audio layers responsible for intra-modal generation; uniform credit assignment, where fine-grained cross-modal alignment regions fail to get efficient exploration. These shortcomings suggest that vanilla RL fine-tuning strategy with a single global advantage often leads to suboptimal results. To address these challenges, we propose OmniNFT, a novel modality-aware online diffusion RL framework with three key innovations: Modality-wise advantage routing, which routes independent per-reward advantages to their respective modality generation branches. Layer-wise gradient surgery, which selectively detaches video-branch gradients on shallow audio layers while retaining those for cross-modal interaction layers. Region-wise loss reweighting, which modulates policy optimization toward critical regions related to audio-video synchronization and fine-grained alignment. Extensive experiments on JavisBench and VBench with the LTX-2 backbone demonstrate that OmniNFT achieve comprehensive improvements in audio and video perceptual quality, cross-modal alignment, and audio-video synchronization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.