acceptodds
Under review as a conference paper at ICLR 2027

The Maternal Phase of Language Model Training

Abstract

What does initialization contribute to the training of a language model, and when and how does training take over from it? Early training is dominated by the initial weights, yet runs that differ only in their seed end with nearly the same loss, so control must pass from what a network is given to what it learns, at a point that has not been located. Embryos make the same handover at the maternal-to-zygotic transition, from materials deposited in the egg to the embryo's own genome, and embryologists locate it by following gene expression through the stages of development. We do the same with a developmental atlas of Pythia that follows the expression of every MLP unit across fifty runs at five sizes and four larger models, and we retrain small models to intervene where the atlas can only correlate. We find that (i) training begins with a maternal phase, in which expression is a copy of initialization set by each unit's projection onto the mean input of its layer; (ii) the phase ends not because the units' weights change, since they keep the initial pattern, but because the mean input rotates, the deepest layers first; (iii) its duration is set by the accumulated update measured against the initial weight scale; and (iv) little of it remains afterwards, and a small network damaged once training is under way restores its loss and its repertoire of unit behaviours but reassigns roles among its units. More broadly, the atlas suggests that the toolkit of developmental biology, from staging to transplantation and lesion, can be turned on training, provided that what a network inherits is first separated from what it learns.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.