Rethinking Audio-Visual Temporal Alignment with Self-Supervised Reinforcement Learning
Abstract
Audio-visual video understanding has become a foundational task in multimodal learning, supporting applications from embodied perception to audio-conditioned video generation, and requires models to jointly reason over temporally aligned audio and visual streams. At the core of this task lies temporal alignment, the binding of audio events to their corresponding visual moments at sub-second granularity. Current omni-modal models often fail at it, a problem commonly referred to as *temporal misalignment* that reflects reliance on coarse audio-visual co-occurrence rather than genuine temporal correspondence. To address this, we propose a **self-supervised reinforcement learning** framework that supervises temporal alignment without external annotations or auxiliary models. We insert paired audio and visual events at controlled onsets into real videos, yielding **AVABench**, an alignment-focused benchmark whose ground truth is defined entirely by the insertion process. The model is then trained with reinforcement learning directly on these construction-derived signals, avoiding the annotation cost of prior supervised pipelines and directly targeting the temporal binding that supervised captions leave unaddressed. Experiments show that existing models achieve only 39.15% accuracy on AVABench, while our 3B **Omni-Aligner** reaches 74.77%, surpassing a 7B baseline by over 20 points, and the learned capability transfers consistently to multiple out-of-domain audio-visual benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.