acceptodds
Under review as a conference paper at ICLR 2027

STAR-AVL: Structured Temporal-Aware Reasoning for Training-Free Video-Centric Audio-Visual Localization

Abstract

Video-centric Audio-Visual Localization (AVL) aims to identify the spatial location of sound-emitting objects by jointly interpreting audio and temporally changing visual information. Existing methods mainly rely on feature-level similarity between audio and visual embeddings, which limits their ability to explicitly reason about the audio-visual context and re-evaluate predictions using temporal information. To address these limitations, we propose STAR-AVL (Structured Temporal-Aware Reasoning for Audio-Visual Localization), a training-free structured reasoning framework built on Multimodal Large Language Models (MLLMs). STAR-AVL consists of two complementary components, Audio-Visual Reasoning and Temporal Context Calibration, which are together organized into four reasoning steps. Audio-Visual Reasoning identifies the actual sound-emitting objects by verifying the audio-visual evidence, while Temporal Context Calibration resolves the remaining localization ambiguity by leveraging the surrounding temporal context. On the AVATAR dataset, STAR-AVL consistently outperforms existing video-centric AVL and CoT reasoning methods. Additional experiments on VGGSound-Duet demonstrate its generalization to image-centric AVL settings. The source code will be publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.