acceptodds
Under review as a conference paper at ICLR 2027

Model Space and Features: Analysis of Internal Representations in Language Models

Abstract

How do model resources shape internal representations, and how does their organization relate to performance? We study feature geometry and superposition in decoder-only Transformers trained on random legal move sequences from Othello and Chess, using rule-defined board-state properties as a common reference. Across 128 Dense and Mixture-of-Experts models varying in width, depth, and FFN structure, increasing width primarily separates already represented state directions, whereas order-dependent state properties depend more on depth and FFN capacity. The relationship between summed feature dimensionality and performance differs across tasks: Chess source-square prediction closely tracks this measure, whereas deeper Othello models improve candidate rankings at similar values. Interventions along estimated feature directions causally change the corresponding output probabilities. Excess loss is approximately log-linear in summed feature dimensionality, and the fitted relationship transfers to FFN structures excluded from fitting. These findings characterize how feature geometry and performance vary with model resources, and locate the performance differences that the measured dimensionality does not account for.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.