acceptodds
Under review as a conference paper at ICLR 2027

Auditing Capacity Trends in Transformer Residual-Stream Geometry

Abstract

Interpreting scaling relationships in transformer residual-stream geometry requires separating measurement conventions, estimation precision, and predictive performance. We audit these factors using released model checkpoints, organizing the analysis around identifiability, reproducibility, trend diagnostics, and internal consistency. Our primary analysis measures late-layer cosine displacement over self-generated TriviaQA answers from six Pythia checkpoints under a prespecified reanalysis protocol. Among constant, capacity-only, and capacity-plus-shape models, the capacity-only model achieves the lowest mean leave-one-model-out absolute log prediction error. With capacity represented by , its fitted exponent is , with a conditional ordinary-least-squares standard error of . Predictive errors remain substantial: the geometric-mean symmetric multiplicative error factor is , reaching in the worst fold. Across three specified analysis windows, the capacity model retains the lowest mean leave-one-out error, although fold-level rankings differ and fitted exponents range from approximately to . Adding a width-to-depth shape term does not improve the mean prediction metric in these comparisons; this finding does not establish the absence of shape effects. Historical extrapolation to two larger Pythia checkpoints further illustrates the limits of predictive accuracy. These results support a family- and window-dependent capacity trend and demonstrate how measurement choices affect its interpretation. Because models generate different text and we do not include fixed-text controls, the reported relationships characterize task-conditioned, self-generated trajectories rather than isolated causal effects of architecture.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.