Query-Oriented and Reasoning-Aware Perception for Versatile Speech Understanding
Abstract
Existing speech understanding systems rely on either implicit speech perception or query-agnostic speech perception, lacking the consideration of the relationship between perception and reasoning. What to perceive depends on the reasoning, while reasoning builds upon perceived information. To address this issue, we propose Query-oriented and Reasoning-aware Perception Optimization (QRPO), a reinforcement learning algorithm that trains a Perceiver to first determine what should be perceived for answering the query and then generate a reasoning-aware description for the perceived information. A Reasoner, which can be instantiated with arbitrary off-the-shelf large language model, performs downstream reasoning based on the description to generate the final response. The correctness of the response is further used as a reward signal to encourage the Perceiver to learn how to describe speech in a way that better supports downstream reasoning. We name the resulting PErceiver-REasoner framework PERE. Experimental results on the MMAU, MMSU, MMAR, and AIR-Bench benchmarks demonstrate that PERE achieves significant gains on speech understanding performance. Furthermore, PERE achieves state-of-the-art results under a more realistic open-ended setting where only the speech and query are provided without candidate answer options.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.