Any Channel Count, One Token Stream: Incorporating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
Abstract
Audio carries phase relations that encode where sources are, yet existing large audio-language models (LALM) reduce it to single-channel content for stability of general audio understanding, leaving essential spatial information for answering questions about direction, distance, and source relations unused. We present SpheraAudio, which accepts recordings with any channel count — mono, stereo, First-Order Ambisonics (FOA), or higher — and maps them into a single token stream. A spatial encoder turns complex multichannel spectra into spatial tokens that are interleaved with audio tokens at corresponding temporal positions, while we propose Spatial Multimodal RoPE (SMRoPE), which makes each token's modality identity and temporal slot visible to rotary attention without altering the language model's original RoPE. Training runs in three stages — spatial-encoder pretraining, projector alignment, and joint supervised-and-reinforcement learning — progressing from geometric supervision to language-based spatial decisions, with answers trained as an explicit spatial chain-of-thought: step-by-step geometric reasoning over the heard scene. Considering that current benchmarks tie their questions to a fixed channel-count format and score answers without spatial reasoning evidence, we introduce SpheraQA, which spans mono, stereo, and FOA recordings and organises questions hierarchically from single-attribute perception through cross-source relations to multi-step reasoning, scored programmatically against simulated scene geometry and real-recording annotations with step-by-step solution references. Across SpheraQA and public spatial-audio benchmarks, SpheraAudio sets a new performance standard for spatial audio understanding. On spatial tasks, it supports high-precision localisation, tracking, and reasoning over static and moving sound sources; beyond them, the model preserves strong general-audio capabilities. We will release the code, the SpheraQA benchmark, and its evaluation metrics to support reproducibility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.