acceptodds
Under review as a conference paper at ICLR 2027

Predictability Is Not Provenance: Rethinking Next-Word Prediction in LLM-Based Neural Encoding

Abstract

Large language model (LLM)-based semantic encoding studies often interpret pre-onset performance in target-aligned lag– curves as evidence that the brain predicts upcoming words. We show that this inference has two limitations. First, contextual representations of successive words are statistically dependent, while responses elicited by those words overlap in the measured signal. Conventional target-word encoding can therefore conflate the target contribution with covariance-weighted contributions from historical sources, making predictive success ambiguous with respect to neural provenance. Second, held-out correlation quantifies aggregate predictive adequacy but does not identify which word generated predictable activity or how its contribution evolves from onset. We formalize this ambiguity with a constrained-availability source-time finite impulse response model (CAST-FIR) and derive Target Isolation from it as a source-adjustment procedure. A strictly causal simulation and a passive acoustic benchmark produce prospective-looking performance without anticipatory computation. Across English ECoG and Chinese SEEG, Target Isolation contracts the cross-onset temporal-generalization area by % and %, respectively. Fitted source-relative kernels peak at ms in English and ms in Chinese after source onset. These results do not rule out neural prediction. They show that target-aligned predictability alone does not establish target-word provenance: performance should establish predictive adequacy, whereas temporal interpretation requires explicit source-time structure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.