When Reward Doesn't See Everything: Modality-Blind Credit Assignment for Simultaneous Speech-To-Speech Translation
Abstract
Reinforcement learning is increasingly used to fine-tune simultaneous speech-to-speech translation (S2ST) models. These models produce both text and audio at the same time, but the reward functions used to train them are computed almost entirely from text. We show that this mismatch causes a failure the training process cannot detect on its own, in which the training reward rises steadily while the model's own voice quality collapses beneath it. We test this by applying group relative policy optimization (GRPO) with a text-only bilingual evaluation understudy (BLEU) reward to Hibiki-Zero, a 3-billion-parameter S2ST model. During training, the reward score rises steadily, but translation quality on held-out data does not follow it, and audio quality falls sharply. Speaker similarity drops by more than 40%, and naturalness, measured by the UTokyo-SaruLab mean opinion score (UTMOS), falls from 3.410 to 2.610. We trace this to GRPO's credit assignment, which broadcasts a single scalar advantage derived from the text reward across every token the model produces, including the audio tokens the reward never examines. This is not simply a missing safeguard: the collapse persists even when we restore GRPO's Kullback-Leibler (KL) divergence penalty, which limits how far the model may drift from its starting point. Holding the reward fixed, we substitute a learned value function that scores each position separately, using proximal policy optimization with generalized advantage estimation (PPO+GAE). This removes the collapse entirely, and it holds when group size is matched to GRPO's. We further test four cheaper alternatives, including one with no learned value estimator at all. All four collapse and are statistically indistinguishable from GRPO. This shows that what matters is not whether a value estimator is present, but whether credit is assigned by comparing samples.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.