Bayesian Speech Synthesizers Can Learn from Multiple Teachers
Abstract
Text-to-speech (TTS) is a "one-to-many" problem: one text can be spoken in many valid ways, so its realizations form a distribution whose spread depends on the input. Most TTS models do not model this spread explicitly: they leave it implicit in the sampling process or fix it in advance. To model this distribution explicitly, we propose **BELLE** (**B**ayesian **e**vidential **l**earning with **l**anguag**e** modelling), which places a Bayesian prior over the mean and variance of each mel-spectrogram frame and predicts this prior from the input, while keeping the architecture and single-pass sampling of common continuous-valued autoregressive TTS models such as MELLE. Estimating this variance requires several recordings of each text, whereas standard TTS corpora provide one. We therefore construct two parallel-speech corpora: **ParallelVox** from real human recordings and **LS+6TTS**, which adds speech from six pre-trained TTS teachers to LibriSpeech (LS). Trained on these corpora, BELLE outperforms MELLE under the same training conditions, is competitive with larger TTS systems on the objective metrics of LibriSpeech test-clean, and supports streaming generation. Analyses of repeated human readings show that real speech has the properties the Bayesian prior assumes and that the variance predicted by BELLE follows the variation across these readings. Parallel speech improves BELLE, and synthetic parallel speech performs comparably to real parallel speech at the same volume and is far easier to scale.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.