Input Embeddings as a Compositional Pronunciation Interface for Frozen LLM-Based Chinese TTS
Abstract
LLM-based text-to-speech systems can misread a character even when its intended pronunciation is known. We study input embeddings as a pronunciation interface for a frozen speech model. A closed-form additive regression maps initials, finals and tones to the embeddings of common characters. To construct a target embedding, we start from real donor rows and add the regression-estimated phonological difference between donor and target. This preserves information in real rows while changing the specified reading; the TTS parameters remain frozen and no recording of the target pronunciation is required. On 120 held-out common-word items, excluding the target's entire initial–final combination in all tones from both regression fitting and donor selection, composition reaches 94% pronunciation accuracy on CosyVoice2 versus 88% for regression alone, with the same ordering on Spark-TTS and GLM-TTS. Under the same exclusion, composition repairs 77% of 193 verified rare-word errors versus 65% for regression alone. In ordinary use, a unified rule that reuses homophones when available repairs 79%, versus 70% for text rewriting. These results connect compositional control to real pronunciation repair, while revealing a trade-off between donor availability and damage to correctly read words.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.