From Speech to Action: Uncertainty-Aware Reinforcement Learning for Tool Calling in Speech LLMs
Abstract
Voice-driven tool calling extends LLM agents from typed interaction to spoken requests, enabling users to invoke external services through natural speech. However, unlike text-based tool calling, speech input introduces an upstream acoustic stage where noise, speaker variation, and ambiguity can distort the request before tool-use decisions are made. Existing uncertainty-aware tool-use methods mainly target text agents, leaving speech-channel uncertainty in tool routing and structured-call generation underexplored. To address this gap, we introduce a hierarchical diagnostic view that separates observable tool-name routing failures from conditional structured-call failures after correct routing, showing that these failure risks can be predicted from Speech LLM representations. This motivates UC-GRPO, an uncertainty-aware post-training and inference framework for end-to-end Speech LLM tool calling. UC-GRPO extracts routing and conditional decisional uncertainty with lightweight probes over Audio Encoder and Decoder states, uses the probe-derived signals to shape GRPO rewards, and implements selective execution through clarification routing. To enable systematic evaluation under controlled acoustic degradation, we further propose SpeechToolBench, a noise-augmented benchmark combining public tool-calling tasks with practical voice-agent scenarios and routing-decision diagnostics. UC-GRPO achieves 51.91% Acc Norm on SpeechToolBench, improving over the base Speech LLM by 11.46 points and standard GRPO by 5.09 points. Zero-shot evaluation on audio-adapted When2Call further shows that the learned uncertainty-aware behavior transfers to an independent tool-use benchmark.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.