Shared-Trajectory Meta-Learning for Joint SFT and RL Optimization in LLMs
Abstract
Supervised fine-tuning (SFT) and reinforcement learning (RL) are two synergistic paradigms for large language models, typically applied in a SFT then RL pipeline. However, the staged schedule can cause RL to overwrite the gains acquired from SFT. Jointly optimizing both is challenging because their gradients may conflict, and existing methods usually consider which signal should dominate while overlooking their interaction. We propose a meta-learning framework that places SFT and RL on a shared inner trajectory and applies a Reptile-style outer update. This implicitly encourages alignment across their gradients, guiding the model towards the regions where progress on one objective is less likely to hinder the other. Our method outperforms SFT, RL, and hybrid methods across multiple backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.