acceptodds
Under review as a conference paper at ICLR 2027

BrainLLM-Bench: Model Rankings Transfer Across Listeners but Not Always Across Stimuli or Readouts

Abstract

Reading a brain-likeness leaderboard as a fact about models assumes its ranking survives a change of readout, listeners or text. On one listener’s MEG, 32 checkpoints get a reliable leaderboard from representational similarity (ranking reliability 0.98) and an equally reliable one from linear encoding (0.96), in opposite orders (ρ = −0.573, family-clustered 95% CI [−0.79, −0.26]). Brain-likeness benchmarks usually fix one readout and resample only participants, so they can see neither this nor whether a ranking survives a change of text. BrainLLM-Bench changes readout, listeners and stimulus one at a time. Listeners barely move a leaderboard: disjoint listener halves agree at median Spearman ρ of +0.76 to +0.96 (two crossed cohorts, both readouts). The reversal, for the readouts as conventionally applied, survives the 48 base checkpoints of the registered test under family-clustered uncertainty (ρ = −0.59, CI [−0.78, −0.12]) but belongs to one cohort, not reproducing on 27-listener MEG or 47-subject English EEG, and feature geometry is not excluded; on the 32 checkpoints it is not determined once static embeddings are also partialled out (ρ = −0.415). Text depends on the cohort: on four narratives heard by the same 27 listeners, similarity leaderboards from different stories do not transfer (mean agreement near zero; exploratory — the registered co-primary supports it, the registered dispersion primary did not), while a fully crossed fMRI cohort (25 narratives × 9 listeners) transfers well (+0.78). The benchmark behind these tests is one protocol over five language cohorts (MEG, EEG and intracranial ECoG; English and German; scalp EEG an uncontrolled probe), six leaderboard models and a readout sweep extended from 32 to 57 checkpoints (48 base) spanning 2019–2026. We prove which feature transformations each readout ignores; this shows the two can disagree without either being wrong, not that this is why they disagree here, and the reversal is a disagreement between measurements, not a refutation of brain-likeness. Model size did not consistently order alignment on a Pythia ladder (70M–2.8B) or across three Qwen generations (up to 72B). A brain-likeness ranking is a claim about a readout, a cohort and a stimulus sample, and should name all three.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.