STAR-AVL: Structured Temporal-Aware Reasoning for Training-Free Video-Centric Audio-Visual Localization
Abstract
Video-centric Audio-Visual Localization (AVL) aims to identify the spatial location of sound-emitting objects by jointly interpreting audio and temporally changing visual information. Existing methods mainly rely on feature-level similarity between audio and visual embeddings, which limits their ability to explicitly reason about the audio-visual context and re-evaluate predictions using temporal information. To address these limitations, we propose STAR-AVL (Structured Temporal-Aware Reasoning for Audio-Visual Localization), a training-free structured reasoning framework built on Multimodal Large Language Models (MLLMs). STAR-AVL consists of two complementary components, Audio-Visual Reasoning and Temporal Context Calibration, which are together organized into four reasoning steps. Audio-Visual Reasoning identifies the actual sound-emitting objects by verifying the audio-visual evidence, while Temporal Context Calibration resolves the remaining localization ambiguity by leveraging the surrounding temporal context. On the AVATAR dataset, STAR-AVL consistently outperforms existing video-centric AVL and CoT reasoning methods. Additional experiments on VGGSound-Duet demonstrate its generalization to image-centric AVL settings. The source code will be publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.