acceptodds
Under review as a conference paper at ICLR 2027

Transformer Depth Up-Scaling: An Infinite-Depth Perspective

Abstract

Progressive training iteratively up-scales small checkpoints to reduce the cost of training large language models. We observe that, as base-model depth increases, layer concatenation leads to greater performance degradation immediately after up-scaling, whereas in-place stacking becomes more stable. Inspired by this depth-dependence trend, we consider an infinite-depth formulation of Transformers, treating depth up-scaling as continuation of parameters and study the consistency problems with a controlled ODE. A criterion on this ODE's terminal-state ensures the consistent next-token predictions between the base and the up-scaled models. For base models of depth , piecewise-constant simple-continuation yields an terminal-state error and thus asymptotically preserves immediate performance. As a concrete continuation scheme, scaled stacking approximates the Euler refinement of simple-continuation, with the same error order uniformly over growth-factors. Moreover, the ODE's adjoint gives -consistent directional derivatives, providing a criterion for preserving descent directions. Benchmarkings on Qwen3 family and progressive trainings on GPT-2 provide supports for these consistency results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.