acceptodds
Under review as a conference paper at ICLR 2027

Empowering Large Audio-Language Models via Spatial Question and Answering

Abstract

Spatial audio understanding requires models to connect sound identities with locations, motion, and relations over time. Most large audio-language models (ALMs), however, process monaural audio and cannot directly access the inter-channel cues available in spatial recordings. We introduce SonicSpatial, a parameter-efficient framework that enhances the spatial understanding capabilities of pretrained ALMs. A question-conditioned spatial selection module uses learnable queries to compress dense features from a frozen spatial audio encoder into compact spatial tokens. These spatial tokens are then combined with the backbone's semantic audio tokens and question tokens as input to the LLM for answer generation. We also introduce SonicSpatial-Bench, comprising 3,045 multiple-choice questions over 576 recorded and simulated FOA audio instances across eight capability categories, including source motion, temporal relations, listener-centered transformations, and room-scale inference. On this benchmark, SonicSpatial with MOSS-Audio-8B achieves 60.5% overall accuracy, improving on its backbone by 22.0 percentage points; the Qwen2.5-Omni variant achieves 63.1%, improving on its backbone by 25.4 points. The two variants reach 64.3% and 65.4%, respectively, on our multiple-choice adaptation of SO-Bench. The spatial selection module can be integrated into diverse ALM backbones without modifying their audio encoders, yielding consistent gains across five evaluated models. These results demonstrate the broad applicability of SonicSpatial as a parameter-efficient spatial adaptation framework.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.