TiCo: Time-Controllable Spoken Dialogue Model
Abstract
Existing spoken dialogue models (SDMs) can generate high-quality responses, yet they struggle to control response duration. Such control is important for managing service costs, improving user experience, and supporting applications with duration requirements. We introduce TiCo-Bench, a benchmark for systematically evaluating duration control across four task categories and target durations of 10–60 seconds. Evaluations of open-source and commercial models reveal frequent failures to meet explicit duration constraints. To address this gap, we propose TiCo, a simple and efficient post-training method that incorporates duration estimation into semantic planning. TiCo uses Spoken Time Markers (STMs), such as <10.6 seconds>, to make estimated elapsed speaking time explicit during autoregressive generation, allowing the model to adjust the remaining content to meet the target duration. TiCo relies on self-generation and reinforcement learning with verifiable rewards, without requiring externally provided question-answer pairs. Across three backbone models, TiCo reduces duration error by 2.7–3.5, and the best TiCo model halves the error of the strongest baseline, while preserving response quality. Demo page: https://tico-9wu.pages.dev/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.