How Does Scaling Reshape Skill Structure in Language Model Representations?
Abstract
Scaling language models consistently improves downstream performance, yet the representation-level changes associated with this improvement remain poorly understood. We study this question through skill structure, using skills as an abstraction of latent capabilities reflected in representations and relevant to downstream prediction. To make skill structure comparable across scales, we develop a task- and template-conditioned analysis based on RSS, with common dimensions and matched random-label calibration. Our theory shows that an increase in the positive calibrated between-skill component raises the guaranteed floor on fitted prototype separation, which in turn can tighten the upper bound on held-out skill-recovery error when prototype-estimation error and within-skill variation are controlled. Across Pythia and Qwen2.5, larger models generally exhibit stronger skill structure. Geometric decomposition shows that this trend is driven primarily by increased separation among skill-conditioned representation centers, accompanied by larger held-out prototype margins. We further test whether this geometry contributes functionally to answer formation. Interventions on skill-aligned representation subspaces produce larger behavioral effects than relative-energy-matched random interventions, supporting functional relevance. The same representation diagnostics provide useful signals for selecting adaptation data under limited training budgets. Together, these results show that scaling is accompanied by stronger skill structure, greater between-skill separation, and more reliable held-out skill recoverability, while the associated skill-aligned subspaces are functionally relevant to answer formation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.