Real-Speech Anchoring for Data-Efficient S2S Post-Training
Abstract
Speech-to-speech (S2S) assistants must adapt to how a request is spoken and reply with appropriate prosody, yet their training supervision rarely carries this information: human-recorded dialogue is scarce and misaligned with assistant behavior, while scalable synthetic dialogue bounds every speech realization by the synthesizer. We call this shortfall the Non-Semantic Information Gap and trace it to the unit both paradigms share, the complete dialogue. At the level of a single exchange, one verified real recording suffices when kept in the role whose realization is the supervision signal, with only its counterpart synthesized. RAST (Real-speech Anchored S2S Training) applies this principle in two directions: user-side anchoring (RAST-U) pairs real input cues with condition-appropriate responses, and assistant-side anchoring (RAST-G) places recorded prosody in the generation targets. We instantiate RAST as 284,660 exchanges spanning 2,914 hours, and introduce RASTBench, a transcript-controlled benchmark that scores response audio directly and separates completing the request from adapting to the speech condition. Against Pure-TTS controls that differ only in the provenance of the anchored endpoint, retaining real speech improves RASTBench by 4.72, 4.52, and 8.81 points for U, G, and their mixture, all statistically significant, with gains concentrated on condition-adaptation criteria. The two directions yield complementary capability profiles, their mixture achieves the strongest response adaptation, and a quarter of the corpus recovers most of the full-data gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.