ReHiGate: Function-Preserving Capacity Augmentation via History-Conditioned Residual Control
Abstract
Large-scale pre-training is expensive, motivating approaches that reuse partially trained models rather than train every target model from scratch. Function-preserving model growth offers a natural solution by expanding the backbone while keeping its function unchanged at insertion. However, forward preservation alone does not guarantee that the added capacity provides useful optimization directions. We propose **ReHiGate** (*Residual History Gating*), which keeps the backbone intact and instead introduces a history-conditioned residual controller. ReHiGate reads a sparse set of preceding-layer states and predicts token- and channel-wise gates for the attention and MLP residual branches, using cross-layer history as control rather than directly mixing it into the representation stream. Its identity-centered initialization exactly recovers the inherited computation while retaining active first-order controller directions, and insertion-time functional diagnostics show that these directions are less aligned with the trained backbone than those introduced by parameter-matched width and depth growth. We evaluate ReHiGate on 124M- and 1.5B-parameter Transformers with continuation endpoints at 10B and 100B tokens, respectively. Across both scales, ReHiGate outperforms plain continuation and parameter-matched width and depth growth at every evaluated insertion point, reducing validation loss by up to 0.044 and 0.017 nats.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.