What Does Model Growth Really Add? Functional Capacity that Survives Continual Learning
Abstract
Model growth adds parameters to a trained checkpoint and continues training, aiming to increase capacity while reusing the existing model instead of training from scratch. But does growing a trained model actually create new functional capacity that survives subsequent training? Parameter count alone cannot answer this question, and “capacity” is itself ambiguous; it can refer to parameter count, effective dimensionality, or the ability to represent new functions. We measure it directly, as the number of new, functionally independent directions a growth step adds, formalized as the rank increment of the model's functional Jacobian and estimated matrix-free at scale. For zero-gated function-preserving growth, this increment is exactly under a transversality condition, i.e. . Continually training a Transformer grown from to parameters across three domains, we find these added directions persist, with all remaining independent after continual learning, even when the model catastrophically forgets (perplexity ) or we deliberately destroy the old function (). Geometric capacity and retained function are therefore decoupled. The functional capacity introduced by growth persists, while the previously learned function is not retained. A second, geometrically distinct mechanism, Net2Net, reaches the same -dimensional capacity from directions that are dormant at initialization () and then retains it, so the decoupling is not specific to the zero-gate parameterization. The implication is precise. Catastrophic forgetting need not reflect a loss of functional capacity, and preserving dimensionality alone is insufficient to preserve the learned function.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.