Equal Depth Exposure, Different Models: Order Effects in Recurrent-Depth Training
Abstract
Recurrent-depth language models apply one shared block a variable number of times, and their training schedules are usually described only by how often each depth is used. We ask whether the order of depths changes the trained model when these counts are held exactly fixed. We train 67M-parameter models on permutations of one fixed multiset of 20,500 per-update depths, with identical counts for each depth across runs. Paired runs share initialization, data order, and learning rates; we evaluate each model at depths 1 to 16. Sorting depths in descending rather than shuffled order raises held-out loss at the deepest trained depth in all five seed pairs, by 0.25 nats on average (95% interval 0.20 to 0.31), and by about 0.6 nats three iterations beyond the trained band. The penalty retains its sign on larger and disjoint evaluation windows, a second multiset, a 102M-parameter model, a C4 corpus slice, and training without either auxiliary loss (three seed pairs). Ascending order, which resembles a depth curriculum, yields losses close to those of shuffled order on depths 5 to 13 but, after training on depths 2 to 10, raises loss by 0.29 nats six iterations beyond the trained band. The penalty depends on when the deepest updates arrive: moving the 5,000 deepest updates from the start to the end of the descending schedule leaves 2 to 6% of the penalty, and 300 to 2,050 further updates at depth 13 bring a checkpoint with 2 nats of excess loss to within 0.03 to 0.11 nats of the shuffled model. Depth order is therefore a training variable in its own right: schedules should be reported by when each depth is used, not only how often.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.