acceptodds
Under review as a conference paper at ICLR 2027

DADC: Dual-Source Anchor Drift Correction for Online Reinforcement Learning

Abstract

Generative policies offer expressive multimodal action representations for online reinforcement learning (RL). However, many existing methods focus on fitting actions favored by the critic, using negative feedback mainly to downweight less promising actions. This provides limited explicit guidance for avoiding low-value actions and can restrict further policy improvement. To address this limitation, we propose Dual-Source Anchor Drift Correction (DADC), an online RL algorithm that combines attractive guidance with explicit repulsion from low-value policy samples. Such guidance requires informative anchors, which are difficult to construct without fixed expert data. DADC therefore constructs attractive and repulsive anchors from distinct sources. For each replay state, we sample initial attractive anchors from an advantage-weighted conditional flow prior trained on replay data and apply several steps of Q-guided Langevin refinement. At the same state, we select current-policy candidates with the lowest critic scores as repulsive anchors, focusing repulsion on relatively low-value actions. The actor update combines direct Q-value maximization with drift guidance toward attractive anchors and away from repulsive anchors. The flow prior and anchor construction are required only during training, while deployment uses a single actor forward pass. Experiments on 14 DMControl and HumanoidBench tasks show that DADC achieves the highest mean return among the evaluated methods on all tasks, delivering substantial gains over the compared generative policy methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.