Does Attention Matter in Motor Imagery EEG Decoders? Diagnosing and Repairing Attention Failure
Abstract
Attention is widely used in hybrid motor-imagery (MI) EEG decoders to dynamically select and integrate evidence for each trial. However, performance gains of the complete decoder and its attention maps do not reveal whether this trial-specific weighting actually contributes to classification. At frozen checkpoints, we replace attention with cross-sample or training-set-mean weights while holding the current trial's Value tensors and all parameters fixed; uniform replacement provides an additional fixed-route control. Across eight decoders on BCIC-IV-2b and BCIC-IV-2a, no architecture shows jointly supported dependence on the two empirical replacements in both datasets. Probes and controlled perturbations trace this failure to token convergence, Value-space equivalence, and downstream bypass or compensation. We therefore develop SOAMNet, which preserves multi-scale temporal trajectories, separates class queries, and forces evidence through one sample-conditioned attention path. Across four MI datasets and three seeds, SOAMNet achieves the highest mean balanced accuracy and Cohen's among 12 baselines. Cross-sample and training-set-mean replacement consistently reduce its balanced accuracy on every dataset; the weaker mean decrease ranges from 18.61 to 40.95 percentage points, and both participant-bootstrap intervals exclude zero throughout. Replacing attention with a matched trainable non-attention head lowers balanced accuracy by 3.51–5.22 points. SOAMNet thus realizes functional reliance on sample-conditioned attention weights while achieving competitive classification performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.