acceptodds
Under review as a conference paper at ICLR 2027

From Task Affinity to Shared Computation: Efficient Spatial Audio Encoding

Abstract

Spatial audio encoders jointly predict sound events, source distances, and directions of arrival. However, assigning a separate processing branch to each task repeats computation even when transformer parameters are shared. We introduce Efficient SpatialAST (ESpAST), a family of encoders guided by layer-wise task-affinity and cross-branch interaction analyses. Representational analysis reveals that sound event detection and distance prediction are most similar in early layers, whereas sound event detection and direction estimation remain comparatively dissimilar. We designed lightweight cross-branch mixers and observed sustained late-layer exchange between branches assigned distance and direction prediction. Together with greater late-layer representational overlap, these learned interactions motivate branch merging, yielding configurations that reduce computation while improving average multi-task performance. On SpatialSoundQA, ESpAST- achieves a average relative improvement across three tasks over a three-branch baseline model, while reducing encoder multiply–accumulate operations by . Evaluation with large language models further shows that reduced-computation encoders retain competitive spatial audio question-answering performance. Our results demonstrate how depth-dependent task relationships guide the allocation of shared and task-specific computation, improving the accuracy–computation trade-off of spatial audio encoders.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.