acceptodds
Under review as a conference paper at ICLR 2027

Timestep-Adaptive Condition Modulation for Flow Matching Text-to-Speech

Abstract

In the field of text-to-speech based on flow matching (FM-based TTS), a vector field (VF) predictor is learned to transport the latent state from noise to a target acoustic representation across multiple timesteps under text and reference speech conditions. Current FM-based TTS methods often concatenate the conditions with the latent state at each timestep and feed it into shared fully connected (FC) layers to obtain a mixed result, which is used to learn the VF predictor for estimating the transport path. For those methods, if the mixing process using FC layers is regarded as the guidance of conditions on the latent state, then this guidance is identical at any timestep. However, the latent state changes dynamically with timesteps. Therefore, unlike previous methods, we suppose that the guidance of conditions on the latent state also changes dynamically with timesteps. To this end, we consider text, reference speech and latent states as conditional distributions with respect to timesteps to optimize the VF predictor for explicitly analyzing whether the guidance changes with timesteps. After deduction, the dependency of the guidance on the timestep is established. Based on this fact, we propose a Timestep-Adaptive Condition Modulation (TACM) block, which uses the timestep embedding to assign two modulation factors for the text and reference speech conditions at each timestep. In this way, the TACM block enables the conditions to adaptively provide appropriate information at different timesteps for guiding the generation of the target acoustic representation. Experimental results show the superiority of our TACM block and validate our hypothesis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.