Pushing Frontier of Lightweight Diffusion for Zero-Shot Text-to-Speech
Abstract
Zero-shot text-to-speech (TTS), which clones an unseen voice from a short prompt, has largely relied on domain-specific priors such as forced alignment or phoneme duration, whose supervision is costly and hard to scale. Recently, large diffusion models trade these priors for scale, letting the alignment emerge from data, while lightweight models for on-device deployment cannot afford this trade. In the absence of a prior, such models fail to learn the alignment reliably, and the explicit priors would reimpose the supervision that scaling was meant to eliminate. In this paper, we present **LiteMel**, a family of mel-spectrogram diffusion models from 0.03B (Small) to 0.13B (Large) parameters that achieves a Pareto-optimal trade-off between model size and performance. Specifically, we first propose a dual-path framework that fuses plain Diffusion Transformers (DiT) favouring intelligibility with hierarchical DiTs favouring speaker fidelity. We also introduce speaker identity that enters through the time embedding and reaches every block by adaptive normalization. Moreover, we propose a learned soft text-speech aligner that restores the monotonic guidance without any supervision. It contains a duration head with a differentiable Gaussian expansion trained by the flow-matching objective alone. To further enhance, we introduce on-policy distillation (OPD) for each student model against our 0.13B member as teacher along the student's own stochastic rollouts, so that the teacher corrects the states the sampler actually visits. Our extensive experiments show that LiteMel-Large attains the SOTA intelligibility, and OPD improves in every metric at both student sizes, cutting the average error by 23% relative to pre-training. The three models together occupy the entire Pareto frontier of average error against parameter count.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.