acceptodds
Under review as a conference paper at ICLR 2027

SyncReward: Human-Aligned Synchronization Reward for Audio-Visual Generation

Abstract

Recent years have witnessed rapid progress in joint audio–visual generation. However, audio–visual synchronization remains a challenge, as generated speech and sound often temporally misalign with visual events. A key bottleneck is the lack of human-annotated data, reliable evaluation benchmarks, and synchronization-aware evaluation models for generated audio–visual outputs. To address these limitations, we conduct a comprehensive study spanning data, benchmarking, and reward modeling. We first introduce SyncReward-Data-80K, a dataset containing 80K professional synchronization ratings for clips from eight audio–visual generators. A taxonomy-guided construction pipeline covers four sound-event families: Speech, Onset, Instrument, and Ambient. We further construct SyncReward-Bench, a held-out benchmark of 1K clips, each independently rated by multiple professionals. Using this benchmark, we systematically evaluate specialized synchronization metrics and omni-models, revealing their substantial misalignment with human judgments. Finally, we develop SyncReward, a reward model with semantic–temporal mixed contrastive Learning on real clips and align with human supervision on generated clips. As an evaluation metric, SyncReward achieves the strongest agreement with human ratings among the evaluated methods. As a reward, it improves synchronization through inference-time best-of- selection and reinforcement learning based post-training. We will release the data, benchmark, model, and code. Additional demos are available at https://anonymous.4open.science/r/SyncReward.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.