From Phonemes to Sentences in the Brain: Composing Neural Primitives for Speech Data Synthesis
Abstract
Compositionality has been proposed as an organizing principle for how neural dynamical primitives are combined to produce movements in simple motor control tasks, such as reaching. Here, we extend this idea to speech by proposing a simple and interpretable mechanistic model of how premotor cortex constructs sentence level dynamics from reusable neural trajectories across multiple hierarchical linguistic levels: phonemes, biphones, syllables, and words. We compare it with a temporal convolutional network (TCN) that can capture non-linear interactions over a longer temporal context. We verify whether the neural activity generated by both models reproduces temporal and spectral properties of the recorded data. Both models preserve speech relevant information: a decoder trained on real neural recordings achieves comparable phoneme decoding performance on synthetic and held out real activity, across all four datasets tested. Finally, we demonstrate that augmenting real datasets with model generated activity improves speech decoding in low data regimes, reducing phoneme error rate by up to 27 percentage points with the TCN and 22 percentage points with the linear model relative to training on real data alone. This is particularly relevant for brain computer interfaces (BCIs), where neural recordings are scarce, while modern decoding models are increasingly data hungry.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.