acceptodds
Under review as a conference paper at ICLR 2027

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

Abstract

Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, whose weights cannot be trained and whose per-call cost makes on-policy reinforcement learning prohibitive, or stay in text: they measure voice agents but cannot improve them. We present SpeechGym, an audio-native agentic environment in which two omni-modal models converse in native audio, with no external ASR or TTS and no API boundary, over the unmodified tasks, tools and success check of an established text agentic benchmark, so that the interaction modality is the only variable and the loop stays local and trainable end to end. The failures speech introduces are perceptual rather than reasoning deficits: the agent picks the right tool and the right argument slot but fills it with a value misheard from the waveform, and that single error cascades into a failed call, a retry of the same call, and a wasted step budget. A second failure is behavioural: under an insistent user the agent performs an unauthorised write and ends the episode believing it helped. Both are trainable, because the environment labels them for free: a call with a misheard argument fails against the database while a correct one succeeds. The obstacle is sparsity, not signal. Outcome-only GRPO is gradient-starved here, since most rollout groups fail identically, while a per-turn process reward crediting each successful tool call restores variance to nearly every group. Trained this way, the agent transfers with no further tuning to an independently implemented voice benchmark, more than doubling task success and lifting an open-weights model above proprietary systems such as GPT-Realtime-2 and Gemini-Live, while using fewer tokens than before training. Along the way it acquires repair behaviours no reward names, such as asking a user to spell a misheard name, and the same checkpoint, which never saw audio question answering, improves on three public audio reasoning and understanding benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.