SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning Paradigms for Large Language Models
Abstract
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models. Our preliminary experiments revealed a phenomenon: SFT suffers from task conflicts under multi-stage training, whereas RL enables coexistence across diverse tasks. Empirically, we trace this to parameter level, observing that RL induces sparse and approximately orthogonal updates across tasks. We provide a theoretical explanation for this mechanism by analyzing multi-task gradient interference. Our results reveal a distinction: interference in SFT is norm-limited, scaling with absolute gradient magnitude, whereas interference in RL is variance-limited, bounded by gradient variance induced by advantage normalization and on-policy optimization. This small variance bound yields near-orthogonal optimization directions across tasks. Leveraging this insight, we propose Parallel-RL, a paradigm that decouples multi-task training, improving efficiency and flexibility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.