Input-Adaptive Acoustic Condition Modulation Across Speech-to-LLM Bridges for Robust Speech Recognition
Abstract
Speech large language models (Speech LLMs) depend on a trainable bridge that compresses acoustic representations into LLM input embeddings. Even with strong frozen speech and language backbones, a bridge learned on one corpus can map unfamiliar acoustic conditions to poorly supported regions of its representation space. We introduce Input-Adaptive Acoustic Condition Modulation (IAACM), a bridge-local calibration method for robust speech recognition. IAACM routes an utterance through a bank of acoustic condition vectors, predicts token-level deviation from a frozen source-bridge manifold, and applies a risk-gated residual FiLM correction. The speech encoder and LLM remain frozen, and inference is self-conditioned on the target utterance without external reference audio. We evaluate matched conditioned and unconditioned systems across eight bridge architectures and six datasets. IAACM lowers the six-domain macro word error rate (WER) for every evaluated bridge and improves 35 of 48 bridge-domain pairs. Averaged uniformly over all pairs, macro WER decreases from 12.22% to 11.65%; the relative reduction grows from 4.68% overall to 17.15% on the highest-error domain of each bridge. Controlled noise and reverberation tests further support robustness. These results demonstrate that bridge-level acoustic calibration improves speech recognition robustness across diverse Speech-to-LLM bridges.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.