FibreTTS: A Fibre-Bundle View on Content-Dependent Control in Text-to-Speech
Abstract
Controllable text-to-speech (TTS) aims to preserve target words while expressing requested attributes such as emotion, speaking rate, and emphasis. Yet language-model-based systems often produce different effects for the same instruction across texts, and stronger control can introduce content errors. To explain this interaction, we propose FibreTTS, drawing on fibre-bundle theory to distinguish content from reading while allowing their relationship to vary. Content representations form the base space, and the possible readings of each representation form its fibre. We construct text- and reference-associated coordinates that explain most of the observed representation variance, providing empirical support for this structured description. Building on this model, we examine two complementary properties and their implications for control. Local triviality provides shared coordinates for what is spoken and how it is read within a content neighborhood, while allowing the same reading change to have different acoustic realizations across texts. Fibre transport, defined through a chosen connection, guides adaptation of a source reading change to target content while retaining the target's own baseline reading. Experiments on two TTS systems, using public emotional speech data and private financial outbound-call data, support the coordinates' behavioral relevance and reveal the effects of direct cross-text control reuse. Transport-inspired residual updates improve the instruction-following scores on outputs whose original control effect was weak, and the improvements are statistically significant on both systems and datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.