How Parameter Correspondence Shapes Cross-layer Optimization
Abstract
Cross-layer optimization couples updates from different Transformer layers and therefore requires a correspondence between their parameters. We introduce bounded depth tempering (BDT), which reshapes AdamW proposals across depth by attenuating high-energy depthwise modes while preserving their common-depth component and pooled centered norm. We compare BDT with correspondence-altered controls matched in common component, centered norm, and angular departure from AdamW at each control branch’s own state. In controlled 124M language models, BDT lowers validation loss relative to both AdamW and these matched controls, showing that correspondence-derived direction carries useful information beyond these measures of update strength. Pretrained layer copying then gives a known relationship between descendants of the same block. Preserving that ancestry retains the benefit, while breaking the ancestry match loses it. Ancestry alone is not sufficient. True descendant pairing outperforms false pairing when update differences are amplified, and the ordering reverses under suppression. We then separate magnitude allocation across ancestry-defined sectors from direction within them. Crossed controls identify magnitude allocation as the more consequential difference. Under a prospectively specified shared window-smoothed direction construction, full-BDT allocations outperform structured-rule allocations, with the ordering reproduced on fresh parents. Relative to fixed ancestry amplification, retained diagnostics show that BDT assigns a larger share of centered proposal energy to descendant differences early in training and a smaller share late. Together, these results establish correspondence, operation, and magnitude allocation as distinct design choices in depth-coupled optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.