acceptodds
Under review as a conference paper at ICLR 2027

Latent Structure, Failed Selection: Representation–Use Gaps in Language Models

Abstract

Performance gaps across language models are commonly attributed to scale, data, and training, yet benchmark accuracy alone cannot distinguish what models represent from how effectively they use those representations. Across 18 base models from six model families, we examined internal structure and forced-choice recall for 5 relation types: ordinal relations, part–whole, category–property, abstract–concept, and attribute–object. Ordinal structure was highly recoverable across model sizes, whereas forced-choice recall increased with size. Erasing ordinal directions selectively impaired matching judgments for temperature, size, quality, and anger after comparison with overlap matched random controls; likelihood showed no detectable effect because baseline recall was near chance in most models. Mapped hidden-state transfer of six novel words spanning temperature, size, and quality outperformed sham injections in 76.8% of 306 ordered sender–receiver pairs. To understand why small models fail, we decomposed answer preferences into question-responsive evidence and question-invariant preference, whose magnitude we call stubbornness. In small models, stubbornness exceeded question-responsive evidence, while signed question-invariant preference was more strongly aligned with lexical priors. To test how question-responsive information reaches the answer, we blocked direct attention from the answer position to the candidate words. This reduced evidence, increased stubbornness, and strengthened its alignment with these baseline word preferences. Because stubbornness is invariant to question polarity, symmetrised scoring cancels this preference without training. Applying the same invariance principle to answer positions improved 126 of 162 model benchmark combinations across nine multiple-choice benchmarks, with gains of up to 23.7 percentage points. The four non-ordinal relation types also showed strong internal structure but only small gains from symmetrisation. Together, these results show that models can represent task-relevant relational structure that fails to control their answers, and that some performance gains with scale arise from using existing representations more effectively rather than from acquiring those representations alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.