acceptodds
Under review as a conference paper at ICLR 2027

ExpertGuard: Transcending the Stability-Plasticity Dilemma in LLM Reinforcement Learning via Expert-Vector Guarded Policy Updates

Abstract

In multi-task scenarios, aligning large language models (LLMs) via reinforcement learning (RL) is governed by the fundamental dilemma between stability—the preservation of performance on historical tasks—and plasticity—the ability to acquire new capabilities. While RL training frequently exhibits greater stability than supervised fine-tuning, it fails to fundamentally resolve gradient conflicts between disparate tasks, thereby yielding suboptimal stability-plasticity trade-offs. Existing solutions, such as dynamic gradient synergy or static model merging, either incur prohibitive computational costs or rely on rigid, non-trainable heuristics that do not align with the dynamic loss landscape during optimization. To address these limitations, we propose ExpertGuard, a novel framework that integrates dynamic expert vector monitoring directly into the RL training loop. Grounded in a rigorous analysis, we demonstrate that the loss of stability arises when new task gradients form obtuse angles with the directions of historical task performance. ExpertGuard projects raw gradient updates onto the orthogonal complement of these dynamic vectors, thereby eliminating components that compromise stability while minimizing perturbation to the primary learning objective to maintain plasticity. Formal proofs and extensive experiments demonstrate that ExpertGuard significantly outperforms existing baselines across different tasks, effectively transcending the stability-plasticity dilemma with negligible computational overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.