acceptodds
Under review as a conference paper at ICLR 2027

RubriVox: From Rule-Based Reward Modeling to Spoken Dialogue Policy Optimization

Abstract

RL (Reinforcement learning) has become the mainstream post-training paradigm for large language models, but its application in end-to-end speech dialogue models is limited by the difficulty of reward modeling. Unlike text-only assistants, spoken dialogue models must be evaluated along both semantic and paralinguistic dimensions. However, existing audio-language reward models often collapse these factors into a single holistic and opaque score, making failures hard to interpret and providing weak learning signals for downstream optimization. We propose RubriVox, a verifier-based RL framework for spoken dialogue models. At its core, RubriVox trains an audio-language Reasoning Reward Model (RubriVox-Verifier) as a structured reward verifier. RubriVox-Verifier decomposes spoken-response evaluation into semantic and paralinguistic critic rules, routes each rule to the proper evidence source for verification, and aggregates the verified rule-level judgments into a scalar reward. This design yields more interpretable and fine-grained feedback than holistic black-box reward scoring. Experiments show that when used for downstream RL, RubriVox provides a principled reward signal for aligning spoken dialogue models with user intent and vocal appropriateness. Demo could be found in https://rubrivox.github.io/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.