CompAdapt: When Audio Representation Adaptation Helps for New Sound Combinations
Abstract
Audio models encounter familiar sounds in combinations absent from downstream training. We introduce CompAdapt, a prototype-guided map that adapts a frozen audio representation. Separate recordings define class-mean prototypes; their sums guide a nonlinear residual map while an anchor keeps it close to the original representation. The fitted map preserves representation width, serves separately trained downstream models, and needs only audio at inference. Across four encoders on ESC-50, it improves mean separation of withheld class combinations by 0.16–0.89 dB at a common 60-epoch endpoint. With each separator’s calibration-selected checkpoint, differences from the original representation range from −0.028 to +0.042 dB. An optional low-rank student retains selected full-map performance within ±0.15 dB on all four encoders with 86.8–96.1% fewer coefficients. On natural SONYC-UST recordings, source-fitted maps improve macro average precision by 0.15–0.19 percentage points in three corrected comparisons. Together with alternative adaptors and cross-dataset reuse, these results show when a reusable representation map helps and how its benefit depends on downstream training. Code and numerical artifacts are available anonymously: https://anonymous.4open.science/r/compadapt-iclr2027-artifact-E83F/README.md
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.