Small Language Models for Large Speech Models: Latent Acoustic Arbitration for ASR Correction
Abstract
Large language models can improve automatic speech recognition (ASR) transcripts, but unrestricted rewriting may trade acoustic consistency for linguistic fluency. We ask whether post-ASR correction requires a text model comparable in scale to the speech model it assists. We propose Latent Acoustic Arbitration (LAA), a parameter-efficient framework in which a compact frozen text model provides linguistic preferences within a bounded speech-generated candidate space, while retained speech-side evidence determines which local changes are accepted. With Granite Speech 8B and a frozen Qwen3-0.6B text refiner, LAA improves the constrained language-selection baseline on all eight evaluated English ASR benchmarks. In a separate scaling analysis spanning refiners from 135M to 32B parameters–a 237x range–WER varies by only 0.14 percentage points. These results provide evidence that effective post-ASR arbitration does not require speech and text models of comparable scale, while controlled 2B and cross-backbone analyses delimit where the observed gains transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.