Acoustic-to-KV Regression for Bioacoustic Recognition in Speech Language Models
Abstract
Animal-sound recognition requires distinguishing call types, callers, and species from acoustic evidence. Speech language models can encode useful distinctions between animal sounds yet fail to select the correct labels. On marmoset calls, a classifier fitted to frozen Qwen2.5-Omni features outperforms audio in-context learning given the same labelled recordings. We introduce Acoustic-to-KV Regression (AKR) to help the model use its acoustic features when forming its own answer. The key idea is to learn from labelled recordings how to adjust the model’s internal states to favor the correct answer. Negative gradients of the correct-answer loss provide correction targets for audio-token key/value (K/V) states. A low-rank ridge regressor then learns to predict these targets from acoustic features. For each new recording, AKR adds the predicted correction to its K/V states before the original decoder generates a label, without backbone updates or query-time backpropagation. Across five recording-group evaluations of 1-kHz marmoset recognition, using support clips that become misclassified after filtering, AKR increases mean accuracy from 12.35% to 32.47% and macro-F1 from 5.92% to 24.54%, outperforming native generation and a fixed correction in every evaluation. Held-out analyses further show that acoustic features predict these correction targets more accurately than fixed-mean and shuffled-pairing controls. These results demonstrate that predicting internal corrections from acoustic features can improve a frozen model’s own recognition.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.