Noise-conditioned Q-Adjoint Anchoring
Abstract
We propose Noise-conditioned Q-Adjoint Anchoring (NQA), an inference-efficient technique that advances the state-of-the-art in offline-to-online reinforcement learning (RL). Expressive RL approaches combining generative policies, distributional critics, and action chunking achieve strong performance, yet incur substantial computational overhead. For efficient inference, multi-step integration in diffusion and flow policies can be replaced with straight flow sampling, but existing approaches encounter two expressivity limitations during training. Specifically, they struggle to (1) directly leverage critic guidance along straight flow trajectories and (2) capture distributional returns while reducing bootstrapping bias. To address these limitations, NQA introduces two algorithmic innovations. First, for policy learning, Q-Adjoint Anchoring applies Adjoint Matching to propagate critic gradients along straight flow trajectories while eliminating multi-step forward SDE integration. Second, for value learning, Noise-conditioned Q-Chunking estimates optimistic values within the return distribution while mitigating bias accumulation and maintaining a short policy execution horizon. In robotic locomotion and manipulation tasks, NQA achieves both high computational efficiency and strong task performance. To the best of our knowledge, NQA is the first method to achieve an average success rate above 90% on one of the most challenging tasks in OGBench, humanoidmaze-giant-navigate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.