Stimulus Overlap Changes Model Comparisons in Cross-Subject Brain Decoding
Abstract
Cross-subject brain decoding holds out the listener, not the story: on MEG-MASC every listener hears the same four stories, so the test story reaches the model twice, through the other listeners in pretraining and, under the standard split, through calibration. A subject-only split does not check this overlap, and the tested model comparison is sensitive to it. Under the standard protocol a dynamic encoder (no residual connections; its pooled feature adds the latent’s start-to-end displacement) beats a matched static convolutional one at low budgets (criterion pre-specified in ×chance at k=100; in rank percentile +0.043 and +0.014, n=10 paired). With story-disjoint calibration (subject split and nominal pretraining volume fixed) the advantage is undetectable; withholding the story from pretraining too, no narrative–budget combination resolves in its favour; at k=2000 the static encoder leads under the standard protocol and with the story withheld. This shows split-sensitivity at the tested sample size, not equal architectures. Transfer survives: withholding the story from pretraining at matched story count (not segment count) leaves about 54% of above-chance signal (k≥500; three narratives, three seeds). Returning 59% of a narrative’s training blocks to a reproduced reference pipeline at matched volume prices the exposure: five exposed runs separate completely from three unexposed (+1.01 to +4.59 top-10 points), and the full dose gives +3.00, same-signed on a second narrative (+2.00), each from a single full-dose run. Still open: whether split or adjacency guard does more, how much of the fully corrected cell (which also withholds a pretraining story) is the split, and how far unequal segment counts move the halved signal. Repeated stories are legitimate for some targets; for claims about unheard stimuli, hold the story out of pretraining and calibration and report a stimulus-derived baseline. We release two stdlib-only tools that flag this overlap; the main results rest on one MEG corpus.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.