acceptodds
Under review as a conference paper at ICLR 2027

Stabilizing Targeted On-Policy Guidance in Agentic Reinforcement Learning

Abstract

Naive on-policy RL is often bottlenecked by the capabilities of the base model, as insufficient coverage can result in an inability to learn from certain rewards. At the same time, rewards in modern post-training systems often come from highly semantic sources, e.g. judge models with a rubric of criteria. Can we utilize such structures to provide a rich directional signal that enables improving on otherwise difficult to optimize rewards? In this work, we develop a framework we refer to as targeted guidance (TGRL) that learns from per-turn textual feedback via an on-policy self-distillation objective. Our focus is not only on directly improving hard-to-optimize behaviors, but on how to do so online without degrading baseline RL performance. Over a range of settings across an agentic multi-turn software engineering environment, we find success by combining targeted self-distillation with a per-criteria RL reward, and by only constructing feedback on turns where the policy specifically fails. We further examine the design choices that lead to stable entropy and reduce oscillatory behavior throughout training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.