acceptodds
Under review as a conference paper at ICLR 2027

Demystifying and Accelerating Emergence to New Tasks in Reinforcement Learning

Abstract

Reinforcement learning with verifiable rewards (RLVR) has shown strong gains in complex reasoning for large language models, but a central question remains whether current post-training techniques truly incentivize models to generalize to new tasks, or whether they mainly reinforce patterns already present in the training data. In this work, we take a mechanistic perspective on this question by investigating a phenomenon called delayed emergence, where a policy first overfits observed rollouts through memorization and much later transitions to out-of-distribution generalization. By analyzing the circuits during training, we show that this capability transition is characterized by a non-monotonic trajectory in the proposed circuit rank and orthogonality index, reflecting a shift from memorization circuits to generalizing circuits. We further identify a critical bottleneck in the RLVR gradient, where low reward variance on training prompts can weaken the signal for transferable features and delay generalization. Based on this insight, we propose Directed Feature Emergence (DFE), a spectral intervention that provably encourages learning on the weak directions in the circuit while preserving the useful representation. Experiments on DELTA-Code, including Manufactoria and BouncingSim, show that DFE consistently accelerates delayed generalization under PPO and GRPO across Qwen3 models, especially in extreme out-of-distribution settings where the base model obtains nearly zero reward. Our results suggest that circuit-level dynamics provide both a diagnostic view of when RLVR moves beyond memorization and a practical target for accelerating task emergence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.