acceptodds
Under review as a conference paper at ICLR 2027

Scale, Drivers, and Limits of Representational Alignment Across Large Language Models

Abstract

Comparative analysis of large language models offers a means of identifying shared representational structure and systematic differences across them. However, aggregate measures of representational similarity, commonly used in these comparisons, can obscure how agreement varies across geometric scales and which linguistic properties support it. We introduce two complementary methods to resolve where models agree and what supports that agreement. Standard representational similarity analysis (RSA) summarizes similarity across all stimulus pairs, potentially obscuring differences in alignment between nearby and distant inputs. Neighborhood-Weighted RSA (NW-RSA) addresses this limitation by weighting stimulus pairs according to their proximity in representational space, enabling systematic comparison of alignment from local neighborhoods to global geometry on a common scale. Counterfactual RSA (CF-RSA) tests which linguistic properties support alignment by comparing intact representations in one model with representations under controlled linguistic perturbations in another. Using these tools across diverse language models, we find strong agreement in coarse representational geometry but substantial divergence in local neighborhoods. Alignment is observed both across broadly sampled natural text and within linguistic equivalence classes that preserve particular aspects of meaning or structure. Counterfactual analyses reveal selective sensitivity, with changes to logical scope reducing alignment substantially more than the tested changes to hierarchical syntax, event roles, or discourse structure. Alignment also decreases when sentence structure or lexical meaning is disrupted through word lists and jabberwocky, suggesting that the regularities of natural language support stronger convergence. These findings provide a differentiated account of representational convergence in language models and refine claims of representational universality by showing that strong overall similarity need not imply agreement in how models represent finer linguistic distinctions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.