Characterizing the Gap Between Knowing and Answering in Speech Language Models
Abstract
A large speech language model can learn a fact from text and answer a written question correctly, but still fail to answer the same question through speech. This discrepancy does not establish whether the learned information is absent from speech processing, is present but not affecting the response, or contributes to a response that still fails to recover the answer. We investigate these possibilities by teaching models new facts through text-only fine-tuning and examining how this learning changes their hidden states when processing speech. By changing the learned answers and selectively removing these changes, we test whether they reflect answer content and influence generation. Our results show that information about the learned answer can remain present and affect the model's response even when speech recall fails. On the speech failures we examine, removing the changes associated with the queried fact harms responses more than removing equally strong changes associated with other facts. We observe this content-sensitive effect across two model families and under changes in question wording and synthesized voice. However, in a controlled ablation study, increasing the measured changes beyond their original strength provides little additional improvement in average answer quality. These findings show that a model can use knowledge learned from text and still answer a spoken question incorrectly. Evidence that this knowledge contributes to a response does not establish that simply strengthening its contribution will produce the correct answer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.