How Far Can Synthetic Speech Take Us? Synthetic Data Recipes for Pretraining Multilingual Speech Language Models
Abstract
Scaling pretraining data has been central to the progress in large language models (LLMs). However, extending scaling recipes to speech-enabled multimodal LLMs is constrained by the scarcity of large, diverse, and permissively licensed speech corpora, a bottleneck even more acute in the speech domain for underrepresented languages. Synthetic speech data can potentially address this gap, but its effectiveness as a training substitute for real data remains unquantified. In this work, we systematically study this question across six languages, varying real-synthetic mixtures, training curricula, and augmentations to the synthetic corpus. Our results show that the quality gap between real and synthetic-only pretraining increases as the training-token budget increases. However, this gap largely disappears when models are pretrained on joint mixtures of synthetic and real speech data: up to 50% of the training data can be synthetic with no detectable degradation across any tested language, while some languages tolerate substantially larger synthetic fractions. Diversifying synthetic data by broadening speaker coverage and performing simple augmentations like reverberation and noise also narrows the synthetic-real gap. Overall, our work provides practical guidance for constructing and using synthetic speech to scale multilingual speech language models under real-world data availability and licensing constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.