ECHO: Embedding Convergence and Hidden Organization in Audio Models
Abstract
Pretrained models trained once on large-scale datasets are now routinely repurposed for entirely new tasks, but there is no reliable way to predict, in advance, how much a given model will benefit from adaptation without first fine-tuning it. This is a real cost in domains where labeled data is scarce, forcing researchers to spend weeks and compute on trial-and-error. We propose ECHO: a concrete, falsifiable geometric signature of adaptation headroom, together with a precise map of the tasks where that signature holds and where it does not. We show that a single measurable property of a model's internal representation; independent of its training data, architecture, or task performance predicts how much a model stands to gain from adapting to a new task, before any adaptation is performed. We study numerous independently trained audio models spanning speech, music, environmental sound, bioacoustics, and machine acoustics. We measure how each model internally arranges sound relative to every other model, and find that agreement between models is real but only partially explained by what they were trained on, no single factor accounts for it fully. On tasks with real learnable signal, models with more clustered representations gain roughly 3.7 times more from adaptation than already-spread models. Directly measuring changes in each model's representational geometry during fine-tuning shows that more collapsed frozen representations tend to reorganize more. This reorganization does not consistently lead to higher accuracy, and the relationship varies across the three adaptation mechanisms. Domain-matched pretraining is therefore not sufficient for model selection: general-purpose models with no domain-specific training outperform domain-specialist models on their own home tasks by up to 7%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.