Adaptive Convolution for Flow Matching-Based Text-to-Speech
Abstract
Flow Matching (FM) is widely used in the field of non-autoregressive text-to-speech (NAR-TTS). For FM-based NAR-TTS, a vector field (VF) needs to be predicted from the current sample state, conditioned on text, speaker information, and the timestep, to transport noise toward a Mel-spectrogram. In the current works, most researchers predict the VF using Transformer, where feed-forward network (FFN) is implemented via 1D convolutional (Conv1D) layers. However, when the training is completed, the kernel weights of Conv1D layers are fixed and cannot adaptively change with different inputs. This leads to high similarity across different frequency bands in the generated spectrograms and the lack of frequency details. To this end, we propose an Adaptive Convolution (AdaConv) block consisting of two components: Cross-Position Kernel Aggregation (CPKA) and Adaptive Gated Kernel Modulation (AGKM). CPKA learns multiple kernel-weight sets with different frequency responses and aggregates them across kernels and positions using Conv1D, thereby increasing kernel weight diversity. As the input evolves across time steps, AGKM computes position-wise gates from the current input to adaptively reweight the contribution of each position. It then extracts global information from the reweighted input and uses it to modulate the aggregated convolutional kernel weights. In this way, the AdaConv block makes the convolutional kernel weights adaptive to the input and helps generate Mel-spectrograms with more frequency details. Then, the AdaConv block is used to implement the FFN in Transformer to predict the VF. Experimental results demonstrate that our method improves both objective and subjective metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.