Aural: Towards Spoken-Content-Aware Spatial Auditory Scene Understanding with Large Audio Language Models
Abstract
Large audio language models (LALMs) have advanced from recognizing *what is present* to understanding *where sources are*. Existing spatial LALMs can model both spoken content and general sound sources; however, these capabilities are largely treated as separate tasks, without explicitly connecting spoken content to spatially grounded sound events. This task separation limits joint understanding of speech and the surrounding auditory scene, where spoken content can be crucial for interpreting the situation. For example, a crash followed by speech may indicate very different situations depending on whether a speaker calls for help or comments on a dropped plate. To address this gap, we first curate **AuralSpeech**, a corpus for spatially localizing and transcribing speakers. We then introduce **AuralScene**, a multiple-choice question-answering dataset for spoken-content understanding, spatial grounding, and relational reasoning. Unlike existing datasets that primarily study these capabilities in isolation, AuralScene explicitly couples spoken content with spatially grounded sound events to support joint scene-level reasoning. To enable its systematic construction, we propose **Spatial Auditory Graph (SAG)**, a structured representation of spatial auditory scenes that enables the generation of diverse and controllable spatial audio scenes and question-answer pairs. Finally, we develop **AuralListener**, a baseline spatial LALM for the proposed tasks that combines monaural and spatial audio token streams, enabling controlled analysis of their complementary roles in spatial auditory understanding. Through comprehensive evaluations, we analyze how spoken content, monaural, and spatial audio representations contribute to understanding and identify remaining challenges in the proposed tasks, providing a step toward **spoken-content-aware spatial audio understanding**.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.