Beyond Turn Timing: Auditing Multi-Axis Rewards for Full-Duplex Speech Reinforcement Learning
Abstract
Full-duplex speech reinforcement learning must align response content and speech quality with the decision to speak, wait, or yield. We present and audit a post-training pipeline that connects generative reward modeling to a streaming speech policy. The design separates deterministic timing rewards from learned semantic and perceptual judgments, preserving distinct feedback about interaction behavior and response quality. Group-relative normalization connects these reward components to policy updates, with sequence-wide or event-local credit assigned through the text stream of joint text–audio rollouts. This provides an explicit interface between multimodal assessment and temporally structured policy optimization. Evaluation on full-duplex interaction and interruption benchmarks examines the integrated pipeline through response coverage, timing, and conditional quality. The recorded configurations exhibit changes in restraint and responsiveness, but do not establish a consistent quality advantage from learned rewards. Response-set analysis further shows why aggregate latency alone can misrepresent these changes. Our study contributes an implemented GRM-to-policy workflow and an empirical account of the capabilities and limitations of multi-axis feedback in full-duplex post-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.