Learning the Optimized Transformer Structure from Generalized Residual Constrained Training
Abstract
Large language models (LLMs) are conventionally trained with a fixed, dense architecture, forcing the entire model to be optimized throughout training despite substantial structural redundancy. Post-training pruning streamlines the model structure for deployment efficiency, but it does not mitigate the training-time costs. In this paper, we propose Generalized Residual Constrained Training (GRCT), a rate–distortion training paradigm built on residual modeling that jointly optimizes task performance and model structure. By optimizing the residual constraints, GRCT directly learns an adaptive model structure during training, reducing training costs as the structure contracts and ultimately yielding a deployable compact model without an extra pruning and recovery stage. In extensive experiments, GRCT outperforms all pruning baselines at the same or higher pruning rates. Compared with the dense model, GRCT even achieves 93.8% of layer pruning while improving accuracy by 1.50 percentage points. Moreover, GRCT speeds up training and inference by up to and , with memory savings of 41.1% and 66.0% on a representative task, substantially improving time and memory efficiency of both training and inference. GRCT shifts the focus of compression research from the final model to the entire training trajectory, providing a unified path toward structure learning for modern LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.