Bounded Integration Windows Gate Hybrid Continuous–Sparse Readouts in a Spike-Converted Transformer Aligned with Language Cortex
Abstract
Large language models and the human brain both turn streams of words into structured meaning. A growing body of work shows that the internal activations of Transformer language models can linearly predict brain responses recorded while people listen to natural speech. How far this convergence reflects shared computation between language models and human brain remains open. Two basic properties of language cortex have been hard to test against models. Cortical regions integrate information over bounded, region-specific timescales rather than indefinitely. Cortical activity also contains sparse, event-like structure, which current model–brain comparisons impose by post-hoc analysis rather than read out from the model. The obstacle is instrumental: standard Transformers expose only continuous activations and a fixed context window. We remove it with a same-weight artificial neural network (ANN) spiking neural network (SNN) substrate. A pretrained Transformer is converted into a spiking network without retraining, keeping the original weights but unfolding computation over discrete timesteps. Continuous, sparse, and hybrid representations can then be read from a single forward pass, and the model-side integration window becomes a tunable scalar. Applied to naturalistic fMRI from 49 listeners, the substrate reveals three coupled phenomena. Brain alignment saturates inside a bounded integration plateau, where the model also shifts toward more compositional layers, a profile that replicates on a second backbone and conversion family. Within this plateau, the hybrid continuous–sparse readout predicts held-out responses better than a matched random control, while a redundant signed variant does not. The size of this advantage is non-monotonic in unfolding time and peaks earlier under faster leak across all twelve language parcels, while the relative hybrid lift descriptively tracks the cortical hierarchy of temporal receptive windows. Together, these results are consistent with a hybrid continuous–sparse code that is useful only inside a finite integration window, and they suggest that bounded temporal dynamics—long studied in spiking systems for biological reasons—may also act as a useful inductive bias for representations that align with the brain.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.