Decouple before Integration: Post-Hoc Synthesis of SFT and RLVR Task Vectors
Abstract
SFT and RLVR improve large language models through complementary learning signals, but existing integration methods rely on training-time choices such as stage lengths, loss weights, or data-mixing ratios. Revisiting this balance after training generally requires another costly optimization run. We therefore ask whether the contributions of SFT and RLVR can instead be calibrated post hoc. Through a task-vector analysis, we find that their updates differ substantially in scale, sign overlap, and module-wise distribution, making direct composition difficult. Motivated by these differences, we propose Decoupled Synthesis (\ours), which independently sparsifies the two task vectors, restores their original norms, and calibrates their relative contributions by searching only two scalar coefficients with Bayesian optimization. We further analyze when such composition can outperform either individually rescaled source. Across seven mathematical reasoning benchmarks, \ours matches or exceeds training-based SFT–RLVR integration methods, while requiring only about of other SFT–RLVR integration baselines' computation when source checkpoints are available. The learned coefficients also transfer to out-of-domain QA and coding benchmarks without re-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.