acceptodds
Under review as a conference paper at ICLR 2027

Real-Speech Anchoring for Data-Efficient S2S Post-Training

Abstract

Speech-to-speech (S2S) assistants must adapt to how a request is spoken and reply with appropriate prosody, yet their training supervision rarely carries this information: human-recorded dialogue is scarce and misaligned with assistant behavior, while scalable synthetic dialogue bounds every speech realization by the synthesizer. We call this shortfall the Non-Semantic Information Gap and trace it to the unit both paradigms share, the complete dialogue. At the level of a single exchange, one verified real recording suffices when kept in the role whose realization is the supervision signal, with only its counterpart synthesized. RAST (Real-speech Anchored S2S Training) applies this principle in two directions: user-side anchoring (RAST-U) pairs real input cues with condition-appropriate responses, and assistant-side anchoring (RAST-G) places recorded prosody in the generation targets. We instantiate RAST as 284,660 exchanges spanning 2,914 hours, and introduce RASTBench, a transcript-controlled benchmark that scores response audio directly and separates completing the request from adapting to the speech condition. Against Pure-TTS controls that differ only in the provenance of the anchored endpoint, retaining real speech improves RASTBench by 4.72, 4.52, and 8.81 points for U, G, and their mixture, all statistically significant, with gains concentrated on condition-adaptation criteria. The two directions yield complementary capability profiles, their mixture achieves the strongest response adaptation, and a quarter of the corpus recovers most of the full-data gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.