Learning Without Listening: Silent Twins for Weak-Evidence Audio-Language Fine-Tuning
Abstract
Post-training of audio language models usually uses supervision with definite answers, yet the available sound may not sufficiently constrain these targets. In this setting, models may still obtain substantial in-domain gains from non-acoustic task signals. Using AVQA fine-tuning as an example, we observe that a large in-domain improvement does not consistently translate into cross-task gains; moreover, replacing all training audio with equal-duration silence still yields substantial in-domain gains. The matched silent twin thus provides a measurable reference for these training-induced changes. In our experiments, we find that normal and silent fine-tuning induce highly overlapping changes across multiple audio language models and transfer benchmarks, and both damage questions that the base model originally answered reliably. We propose STS (Silent-Twin Subtraction): it uses the weight update produced by silent fine-tuning as a reference, subtracts a tunable proportion of it from the normal fine-tuning update, and uses a second coefficient to control the overall magnitude of the combined update. Across models and benchmarks, silent twins reveal what fine-tuning learns without listening, and STS consistently outperforms conventional fine-tuning on every transfer benchmark we evaluate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.