MPO-PRO: TRAJECTORY SHAPING THROUGH TEMPO- RARY MATRIX FACTORIZATION
Abstract
Post-training adapts pretrained language models to tasks such as mathematical reasoning and code generation, but local deployment often limits model size and inference cost. One way to improve post-training is to use overparameterization during training, allowing a richer parameterization than the final deployed model. We propose MPO-PRO, an overparameterization-based approach that enriches post-training by using matrix product operators (MPOs) as temporary training coordinates for dense updates and LoRA factors, and contracts them before deployment. This keeps the original model architecture or adapter rank unchanged while providing a richer training parameterization. For the fixed-rank branch, we use the first-order LoRA-Pro update as a reference and approximately match it in the temporary MPO representation.On GSM8K and HumanEval, the dense and fixed-rank branches outperform full fine-tuning and LoRA-Pro, respectively. When trained on SciInstruct, MPO-PRO also outperforms LoRA-Pro on SciBench Physics. We further analyze the update mechanism through decomposition and find that the performance gain is mainly associated with the first-order implementation deviation introduced by the temporary representation relative to LoRA-Pro.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.