Synthesizing Speech from Intracranial Recordings with Conditional Flow Matching
Abstract
The task of synthesizing speech from neurological data has advanced significantly over the last five years, through a combination of higher-quality data and modern machine-learning techniques. To date, all such approaches have been based on discriminative modeling, i.e. building a deterministic map from neural data to speech audio (e.g., the spectrogram). However, there is reason to believe such models are unnecessarily restrictive. Inspired by the recent evolution in text-to-speech from discriminative to generative models, we reinterpret speech synthesis as a conditional-generation problem. We use conditional flow matching to learn a velocity field that transports a standard normal distribution to the distribution of log-mel spectrogram features, with velocity predictions conditioned on the corresponding neural data by way of a convolutional encoder. Discriminative modeling emerges as the special case in which the model is trained to predict only the initial velocity. We demonstrate our approach on (1) a set of stereo-EEG/speech audio recordings recently made by our group from two participants, and (2) a public data set of microelectrode recordings (single units) made from a person with severe dysarthria. In both cases, our conditional flow outperforms all competitors, in the latter case bringing word error rates of (transcribed) synthesized speech down from about 56% to 22%. Our model also requires about one fourth as many training trials to achieve comparable intelligibility. Overall, we show that conditional flow matching is a powerful approach to neural speech synthesis for datasets ranging from small (sEEG, 0.5 hours) to large (single units, 38 hours).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.