acceptodds
Under review as a conference paper at ICLR 2027

One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control

Abstract

Although Reinforcement Learning with Verifiable Rewards (RLVR) provides a principled post-training paradigm for cultivating complex reasoning abilities of Large Language Models (LLMs), jointly training one foundation model across heterogeneous domains usually introduces multi-domain trade-offs and destabilizes the overall policy optimization. While existing approaches typically diagnose cross-domain interference via gradient alignment or curvature proxies, they usually overlook interference that accumulates across consecutive updates—where seemingly non-conflicting gradients can still overwrite prior policy gains. We show that these cross-step dynamics are directly tractable between adjacent checkpoints through token log-probability footprints, bypassing costly curvature approximations. Building on this observation, we propose OSOL, which organizes each mixed-domain update around a single focus domain and uses the preceding checkpoint footprint to identify rebound-prone tokens and apply a rank-based, adaptively scaled correction. Our analysis shows that this correction contracts the targeted negative-to-positive backtracking component, while controlled studies show that cross-step backtracking is more strongly associated with subsequent task damage than within-step gradient diagnostics and that checkpoint footprints identify future rebound more effectively than Hessian-based proxies. Across Qwen3-8B-Base and Qwen3-30B-A3B, OSOL achieves the best four-domain macro average on both backbones, while ranking first or second in most evaluations. On Qwen3-30B-A3B, it reaches a domain-macro average of 0.4822, outperforming the strongest baseline by 5.7%. Ablations support prioritizing one domain per step, using checkpoint drift to locate rebound-prone tokens, and correcting them at the token level. Code is available at https://anonymous.4open.science/r/anonymous-osol-83DF/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.