acceptodds
Under review as a conference paper at ICLR 2027

Role-Dependent Use of Mixed-Quality Data in Offline RL

Abstract

Offline reinforcement learning learns from fixed datasets whose transitions may vary substantially in behavioral quality. Data generated by poor behavior, how- ever, may still provide useful transitions for critic learning or useful states for policy improvement even when their recorded actions are undesirable targets for imitation. We study this distinction by separating three roles of offline data: critic learning, Q-guided policy improvement at recorded states, and behavioral super- vision from recorded actions. Controlled routing experiments with expert–random mixtures show that the effect of the same additional data can differ substantially depending on the learning operation in which it is used, and can even change sign across environments and data regimes. Motivated by these findings, we in- troduce Trust-Mix, which uses the full dataset for critic learning and Q-guided policy improvement while selectively weighting recorded actions for behavioral supervision, without requiring source-policy labels. Experiments across three continuous-control environments and four expert-data budgets show that the bene- fit of additional random data varies across environments, data regimes, and offline RL algorithms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.