The Two Clocks Of Representation Learning In Language Models
Abstract
Representation learning runs on two clocks. The token clock, which tokens a feature answers to, is largely done within the first percent of pretraining. The wiring clock, how it answers within those tokens, calibrates to the end of the run, in step with the direction the feature writes. We track sparse-autoencoder features across checkpoints of three Pythia scales and OLMo-2-1B, splitting each activation profile exactly into token-set and context parts. By step 1k the token set reads r = 0.71–0.78 against its final form, while context selectivity and direction stand at 0.22–0.42 and move together to the end. The window the two clocks open is where a feature’s mature role becomes readable. Containment, the nesting of one feature’s firing set inside another’s, ranks at step 1k which persisting features will organise the mature network, at AUC 0.74–0.89 with nothing fitted, while which features persist stays near chance. Those persistent, high-containment features are the model’s carriers. They cost 1.7–1.9× as much to remove as populations matched on layer, breadth and activation energy, and ablating a containment parent cuts its child’s firing by 11% at the two smaller Pythia scales.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.