acceptodds
Under review as a conference paper at ICLR 2027

FibreTTS: A Fibre-Bundle View on Content-Dependent Control in Text-to-Speech

Abstract

Controllable text-to-speech (TTS) aims to preserve target words while expressing requested attributes such as emotion, speaking rate, and emphasis. Yet language-model-based systems often produce different effects for the same instruction across texts, and stronger control can introduce content errors. To explain this interaction, we propose FibreTTS, drawing on fibre-bundle theory to distinguish content from reading while allowing their relationship to vary. Content representations form the base space, and the possible readings of each representation form its fibre. We construct text- and reference-associated coordinates that explain most of the observed representation variance, providing empirical support for this structured description. Building on this model, we examine two complementary properties and their implications for control. Local triviality provides shared coordinates for what is spoken and how it is read within a content neighborhood, while allowing the same reading change to have different acoustic realizations across texts. Fibre transport, defined through a chosen connection, guides adaptation of a source reading change to target content while retaining the target's own baseline reading. Experiments on two TTS systems, using public emotional speech data and private financial outbound-call data, support the coordinates' behavioral relevance and reveal the effects of direct cross-text control reuse. Transport-inspired residual updates improve the instruction-following scores on outputs whose original control effect was weak, and the improvements are statistically significant on both systems and datasets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.