acceptodds
Under review as a conference paper at ICLR 2027

Computational Dissection of Sound-to-Meaning Hierarchies in Speech Foundation Models and the Human Brain

Abstract

Whether speech foundation models and the human brain abstract sound into meaning along a convergent route remains unresolved: prior comparisons rely on brain scores, which are hampered by low-level confounds and do not address compositional meaning. We introduce a spoken multi-definition-to-target paradigm in which acoustic, phonetic, and lexical form vary while meaning converges: differently worded definitions, read by multiple speakers, converge on a targeted semantic concept. We dissect eleven speech foundation models spanning masked prediction (HuBERT, WavLM), contrastive prediction (wav2vec 2.0), weak supervision (Whisper), and audio-LLM encoders (Qwen2-Audio, Qwen3-Omni), as well as intracranial recordings from 29 patients performing the same task. Layer-wise decoding and representational geometry analyses in the models reveal cascaded dynamics: speaker identity separates first, followed by surface form and phonetics, each persisting across most layers, whereas lexico-syntactic category and target semantics emerge last, in narrow bands whose depth is set by the training objective, and the semantic target is decodable from held-out definitions. Temporal response function encoding with interpretable speech and linguistic features shows that continuous acoustic and prosodic features trade off against categorical linguistic features across depth, whereas cortex abstracts along a graded, nested hierarchy from core auditory regions to distributed frontal, parietal, and medial hubs, in which part of speech (POS) behaves as a lexical property and constituency structure engages posterior temporal and parietal cortex. Temporal generalization decoding shows that the two systems share dynamics for phonetic and statistical word-level information but diverge for rule-based structure and sentence-level semantic categories. Low-level information is thus spread across depth in the models but localized in cortex, and higher-order structure shows the reverse pattern. By testing speech models for brain-like abstraction rather than brain-like prediction, this work establishes a framework that locates where and how artificial and biological speech networks converge and diverge, and points to the model layers that could serve neuroprostheses decoding meaning as well as speech.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.