Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning
Abstract
Critic-free group-based RL has become a scalable paradigm for LLM post-training. However, its effectiveness is constrained by a major limitation: rollouts are allocated uniformly even though their learning value varies substantially across tasks and trajectory states. Despite recent efforts to allocate rollouts adaptively across tasks, existing methods suffer from two core gaps: the non-stationary gap, which calls for intervention strategies from learning an evolving policy rather than relying on heuristic rules, and the non-scalar gap, which calls for structured control over where and how to intervene rather than merely deciding how many rollouts to generate. To close these gaps, we introduce Recoverability-Aware Intervention Learning (RAIL), a training-time framework that turns rollout generation into a learnable intervention process by optimizing structured decisions according to their realized recoverability gains. Specifically, RAIL first casts intervention selection as an online contextual-bandit problem and then trains a recoverability controller from intervention traces collected through a shadow-to-live procedure, performing intervention learning while policy evolves. Finally, we evaluate RAIL for its effectiveness, adaptivity, expressiveness, and efficiency, showing consistent gains with constrained rollout budgets. Together, our results establish recoverability-aware intervention as a principled path toward more informative rollout generation, enabling post-training to optimize from stronger and less redundant learning signals.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.