acceptodds
Under review as a conference paper at ICLR 2027

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

Abstract

Recent years have witnessed remarkable progress in joint audio-video generation. Nevertheless, existing models still suffer from limited per-modality fidelity, inadequate text-modality alignment, and weak cross-modal synchronization. Reinforcement-learning-based post-training offers a promising avenue for addressing these shortcomings. However, naively extending such approaches to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle the learning signals of the two modalities and obscure credit assignment. Jointly optimizing both modality towers also incurs substantial computational cost despite their distinct optimization dynamics. Furthermore, the difficulty of synchronization evaluation varies with the sampled counter-part modality, making fair reward comparison difficult. To address these challenges, we propose AV-GRPO, a modality-anchored online diffusion reinforcement learning framework, together with 5DAV, a fully decoupled and difficulty-controllable training dataset. AV-GRPO integrates three key components: (1) modality-anchored rollouts that disentangle learning signals while reducing anchor-induced difficulty variation; (2) trajectory-locked frozen-tower optimization that reduces training costs and redirects credit assignment, thereby simplifying optimization. and (3) adaptive objectives and perturbation strengths tailored to modality-specific dynamics and sample difficulty. Collectively, these designs transform coupled multi-modal preference learning into a set of conditional unimodal subproblems, enabling more accurate reward attribution and more effective synchronization optimization. We construct 5DAV, a dataset decoupled along five dimensions, to facilitate systematic training. Experiments on JavisBench and VABench demonstrate that AV-GRPO consistently outperforms LTX-2.3 in generation quality, semantic alignment, and cross-modal synchronization under both LoRA and full fine-tuning settings. Extensive ablation studies further validate the effectiveness of the proposed designs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.