Token-Level Trust Regions under Reuse: Component Effects, Stability, and Intervention
Abstract
Token-level trust regions constrain policy updates in reinforcement learning of large language models, but local constraints do not by themselves explain accuracy gains or stability under repeated use of generated answers. We study these questions in TROLL through component comparisons, update-budget controls and same-checkpoint interventions. On Qwen3-1.7B/GSM8K, removing gradients through an active projection causes no detectable accuracy loss, and at low reuse unclipped policy gradient recovers most of the gain over default-range clipping. At the published radius, two passes over each batch collapse in all three tested seeds, reaching 55.47% mean final accuracy, whereas allocating the same number of optimizer updates to twice as many fresh answers keeps all three runs healthy and reaches 86.83%. Smaller radii reduce this instability. From three shared alarm checkpoints, removing the second pass improves final GSM8K accuracy by 28.20–47.08 points over continuing two passes and by 3.11–3.64 points over the starting checkpoints. These interventions establish the value of changing reuse at those states, separately from choosing when to act. To explore when to intervene, we formulate Co-Drift Guard (CDG), a candidate label-free rule based on joint growth of entropy and within-step pre-projection KL, and separately evaluate model retention, reuse switching and detection across configurations. The results show that the benefit of removing the second pass does not automatically confer an advantage on alarm-based timing, and that detection performance depends on the training configuration. Our study distinguishes component-level gains, stability under reuse and intervention effects, providing empirical foundations for further study of state-dependent reuse decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.