Evaluating OOD Actions in Offline RL via Quantile Advantage
Abstract
Offline reinforcement learning must improve on recorded behavior while controlling errors in evaluating out-of-distribution (OOD) actions. We introduce Advantage-Based Diffusion Actor-Critic (ADAC), which compares candidate and behavior actions through their successor values. A learned dynamics model predicts these successors; a state-value function learned from offline transitions evaluates them. Subtracting a quantile of behavior-induced values from the candidate's successor value defines its quantile advantage. This signed signal corrects critic targets, guiding a diffusion actor alongside behavior cloning. The comparison covers in-distribution and OOD actions; its accuracy depends on the learned dynamics and values. We characterize reference-value updates and the augmented fixed-policy evaluation operator. ADAC achieves the highest reported mean on 15 of 20 D4RL tasks. Four-domain component ablations support the correction's contribution, and a 15-task quantile study examines its reference threshold. PointMaze illustrates shorter routes beyond the recorded patterns. Code: https://anonymous.4open.science/r/adac-14D0.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.