acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Collapse in Mixed-Policy RL: The Underlying Mechanism and A Unified Resolution Framework

Abstract

Compared to pure on-policy reinforcement learning (RL), mixed-policy RL integrates the exploration of online rollouts with the exploitation of off-policy external traces, effectively surpassing the inherent capability boundaries of the base large language model. Nevertheless, directly applying standard importance sampling to integrate external data induces performance degradation. Although prior methods circumvent this instability by redesigning the surrogate objective, the fundamental cause of the collapse and the unifying theoretical mechanism underlying these disparate solutions remain poorly understood. In this paper, we reveal the fundamental cause of training collapse in mixed-policy RL. Our analysis demonstrates that the instability is not explained by the variance of importance ratios, but driven by severe gradient conflicts induced by low-probability external tokens against the rest of the batch. Motivated by this insight, we propose a unified gradient-suppression framework that reveals existing solutions, despite their different formulations, are fundamentally united by an implicit mechanism that suppresses gradient weights of low-probability tokens. To demonstrate the validity and utility of this framework, we derive a direct probability-aware reweighting method as a concrete instantiation. Extensive evaluations across multiple backbones and benchmarks demonstrate the superior performance of this method, while simultaneously validating the theoretical soundness of our framework and its predictive capacity for method design.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.