acceptodds
Under review as a conference paper at ICLR 2027

LENS: Linking Phoneme, Word, and Sentence Supervision for Brain-to-Text Decoding

Abstract

Brain–computer interfaces offer a path to restoring communication by decoding intended speech from neural activity. High-performing brain-to-text systems combine connectionist temporal classification (CTC)-based phoneme decoding with language models, yet standard CTC neither explicitly encodes phoneme relations nor directly supervises word- and sentence-level neural representations. Pretrained linguistic representations could provide that direct supervision, but connecting words to neural features is challenging without annotated word timings. We propose **LENS**, a framework that **l**inks phon**e**me-, word-, and se**n**tence-level **s**upervision for brain-to-text decoding. LENS combines classifier geometry with neural representation covariance to assign bounded credit to phoneme alternatives during CTC supervision. Building on phoneme-level training, we introduce Posterior-Guided Multiscale Linguistic Alignment: model-derived soft alignment couples neural features with pretrained word embeddings without annotated word timings, complemented by utterance-level alignment with pretrained sentence representations. LENS achieves state-of-the-art word error rates on the Brain-to-Text ’24 and ’25 benchmarks without changing the established decoding pipeline. Together, these findings support combining phoneme supervision informed by neural evidence and phoneme relations with word- and sentence-level linguistic guidance for more accurate brain-to-text decoding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.