acceptodds
Under review as a conference paper at ICLR 2027

The Two Clocks Of Representation Learning In Language Models

Abstract

Representation learning runs on two clocks. The token clock, which tokens a feature answers to, is largely done within the first percent of pretraining. The wiring clock, how it answers within those tokens, calibrates to the end of the run, in step with the direction the feature writes. We track sparse-autoencoder features across checkpoints of three Pythia scales and OLMo-2-1B, splitting each activation profile exactly into token-set and context parts. By step 1k the token set reads r = 0.71–0.78 against its final form, while context selectivity and direction stand at 0.22–0.42 and move together to the end. The window the two clocks open is where a feature’s mature role becomes readable. Containment, the nesting of one feature’s firing set inside another’s, ranks at step 1k which persisting features will organise the mature network, at AUC 0.74–0.89 with nothing fitted, while which features persist stays near chance. Those persistent, high-containment features are the model’s carriers. They cost 1.7–1.9× as much to remove as populations matched on layer, breadth and activation energy, and ablating a containment parent cuts its child’s firing by 11% at the two smaller Pythia scales.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.