Learning When to Think While Listening in Large Audio-Language Models
Abstract
Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reasoning quality and responsiveness are tightly coupled: reasoning that starts after the speech endpoint remains on the response path, while reasoning under partial audio risks missing decisive late evidence. We introduce a learnable wait-think-answer control formulation for LALMs. Motivated by the incremental nature of human conversation, the controller decides under partial audio evidence when to wait, and when to produce a compact reasoning update. Using Qwen2.5-Omni-7B as the base model, we construct aligned wait-think-answer traces from spoken reasoning data, train the controller with supervised fine-tuning (SFT), and then apply Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO). The reward combines answer correctness, action validity, update timing, post-end residual work, reasoning quality, and chain consistency, optimizing the complete wait-think-answer trajectory and not the final answer alone. On a six-task synthetic spoken reasoning question answering (SRQA) benchmark, the six-reward DAPO controller improves the row-weighted accuracy from 67.6% to 70.3% while reducing post-endpoint residual reasoning from 10.44 to 8.99 tokens under the same Qwen streaming replay protocol. These tokens measure remaining explicit reasoning, not elapsed response time. Native-direct accuracy is essentially unchanged after controller post-training. On a 186-item human-recorded Real Audio Bench, a transfer check beyond text-to-speech (TTS)-rendered speech, the controller family remains functional: SFT achieves the strongest accuracy, while the six-reward DAPO controller is the only learned variant whose post-end residual falls below the base. These results show that, in the endpoint-gated Qwen setting, a streaming model can learn when to make intermediate reasoning explicit during the audio stream.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.