UNDERSTANDING AND IMPROVING CROSS-SESSION GENERALIZATION IN SURFACE-EMG SILENT SPEECH RECOGNITION
Abstract
Silent-speech decoding from surface electromyography (sEMG) offers a non-invasive communication interface, but performance can deteriorate across recording sessions. We use Mindscape-137, to our knowledge the largest open silent speech sEMG corpus by utterance count, with 25,957 silent utterances across 25 sessions from one speaker spanning 21,632 unique sentences and 2,951 word types, to examine signal variation and evaluate interventions for cross-session recognition. Repeated sentences and electrode reapplication between sessions support content-matched comparisons. Channel amplitudes and relationships vary across sessions, while temporal amplitude-envelope features remain informative about repeated-prompt identity. Imposed gain and mixing perturbations increase decoding errors, so the decoder is sensitive to these transformations. Out-of-span augmentation reduces word error rate from 72.15% to 65.77%. In a session-scaling study with three fixed held-out sessions, increasing training coverage from 5 to 22 sessions reduces fresh-sentence word error rate from 67.22% to 28.30%. Checkpoint averaging and latent covariance alignment provide further improvements in the scaling study, whereas the tested rate correction and labeled calibration do not consistently help.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.