Breaking the Curse of Repulsion in Off-Policy Reinforcement Learning
Abstract
Off-policy policy optimization reuses historical behavior, including negative- advantage samples that provide useful corrective feedback. We identify learner- relative remoteness as a key state variable of reuse: repeated updates from a fixed negative sample can make it increasingly remote without weakening its score re- sponse. This creates a self-reinforcing remoteness–reuse loop. Aggregate analysis shows how this feedback competes with positive attraction, ranging from stable displacement beyond the positive-only solution to persistent drift and loss of finite stable equilibria. The framework covers both continuous Gaussian and discrete categorical policies, with distinct far-field manifestations: Gaussian score response grows with standardized distance, whereas categorical score response remains bounded while suppression persists. Dynamic Remoteness-Aware Policy Optimiza- tion (DRPO) exponentially attenuates remote negative updates while preserving near-field feedback, with Gaussian ultimate-boundedness guarantees. Controlled experiments separately isolate policy geometry as the source of the far-field effect and establish its causal transmission to instability. External diagnostics in D4RL locomotion and structured generation recover the predicted near-to-far pattern, while task-level evaluations show that DRPO achieves higher aggregate scores than reported matched controls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.