acceptodds
Under review as a conference paper at ICLR 2027

Beyond Turn Timing: Auditing Multi-Axis Rewards for Full-Duplex Speech Reinforcement Learning

Abstract

Full-duplex speech reinforcement learning must align response content and speech quality with the decision to speak, wait, or yield. We present and audit a post-training pipeline that connects generative reward modeling to a streaming speech policy. The design separates deterministic timing rewards from learned semantic and perceptual judgments, preserving distinct feedback about interaction behavior and response quality. Group-relative normalization connects these reward components to policy updates, with sequence-wide or event-local credit assigned through the text stream of joint text–audio rollouts. This provides an explicit interface between multimodal assessment and temporally structured policy optimization. Evaluation on full-duplex interaction and interruption benchmarks examines the integrated pipeline through response coverage, timing, and conditional quality. The recorded configurations exhibit changes in restraint and responsiveness, but do not establish a consistent quality advantage from learned rewards. Response-set analysis further shows why aggregate latency alone can misrepresent these changes. Our study contributes an implemented GRM-to-policy workflow and an empirical account of the capabilities and limitations of multi-axis feedback in full-duplex post-training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.