MOST-OPSD: Dual-Level Speech-Prior Distillation for High-Performance Language BCIs
Abstract
Language brain-computer interfaces (BCIs) offer promising approaches to restoring communication, but accurate speech decoding remains challenging with limited neural recordings. Existing methods mainly improve neural decoding by using phoneme context and enhancing temporal modeling. However, decoding models struggle to learn a reasonable continuous speech structure from discrete phoneme label sequences with transcript supervision alone, thereby limiting decoding performance and generalization. Here, we propose the **MOST-OPSD** neural decoding framework, combining Mixture of Speech Teachers with privileged On-Policy Self-Distillation, which transfers transcript-generated speech priors to the decoding model at the representation and prediction levels. Specifically, the MOST module transfers structured speech representations across multiple neural stages via independent projections and length-aware Soft-DTW loss, while the OPSD module distills speech-conditioned phoneme distributions by querying speech features using current neural states and applying a shared phoneme head. On the *Brain-to-Text Benchmark '24*, the complete system of the MOST-OPSD neural decoding framework with the language model achieved state-of-the-art performance, achieving a word error rate (WER) of **3.63%** on the hidden test set. Supplementary experiments on the second dataset *Brain-to-Text '25* also show that our method significantly outperforms baselines. The results indicate that the proposed dual-level speech-prior distillation framework provides a promising technical path for high-performance language BCIs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.