SONATINE: Spatial Audio Understanding via Tool-Integrated Reasoning in Audio LLMs
Abstract
Most general-purpose Audio Large Language Models (Audio LLMs) accept only single-channel audio and therefore cannot use inter-channel spatial cues. Existing approaches to spatial audio understanding internalize these cues by feeding dense features extracted from multichannel First-Order Ambisonics (FOA) recordings into the model. However, such adaptation improves spatial understanding at a potential cost to general audio understanding. Spatial measurement can instead be externalized: signal-processing estimators compute quantities such as sound direction directly from FOA recordings, and their outputs are provided to an unchanged Audio LLM. But to reliably measure the queried target, such an estimator must operate on a target-informative window in which the queried sound is clear and overlap is limited. Because the estimator measures all active sounds in its input window, spatial information from other sources can otherwise be mixed into the measurement. To address this problem, we introduce SONATINE, a tool-integrated reasoning framework in which spatial estimators act as tools and the Audio LLM specifies a time window as an argument to each tool call. The Audio LLM, with no added spatial module, receives the question and only the mono channel, while the selected tool accesses all FOA channels within the specified window and returns a spatial observation. Within SONATINE, obtaining reliable spatial measurements requires adaptive window selection, in which the Audio LLM uses the question and mono audio to choose a target-informative window for each tool call. To teach this behavior and general spatial tool use, we construct SONATINE-15k, a dataset of 14,990 tool-use trajectories, and use it for Supervised Fine-Tuning (SFT), followed by Group Relative Policy Optimization (GRPO). On SO-Bench, our 7B model surpasses all evaluated internalized spatial baselines while using far fewer spatial QA examples, with its largest advantage under heavy source overlap. By externalizing spatial measurement, SONATINE also enables spatial audio understanding in closed-source models through prompting alone, helps retain general audio understanding, and supports transfer from FOA to a six-channel format without further training by changing only the tool implementations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.