PSIL: Learning Safe Offline Policy from Preferred and Non-Preferred Demonstrations
Abstract
Offline imitation learning (IL) typically relies on expert (i.e., preferred) demonstrations and a large collection of supplementary trajectories for broader state-action coverage. However, existing approaches often struggle to effectively integrate the valuable signal from non-preferred demonstrations, which provide critical information about what the agent should avoid. We propose PSIL, Progressive Safe Imitation Learning, a self-training-based offline IL algorithm that can flexibly learn from preferred demonstrations, non-preferred demonstrations, or both. PSIL progressively propagates behavioral information from this small set of demonstrations to a large supplementary trajectory dataset by assigning preference-weights based on their similarity to preferred or non-preferred behaviors. These weights induce a relative ordering over trajectories and are used in weighted behavior cloning to emphasize preferred behavior while down-weighting non-preferred behavior. We evaluate PSIL in Learning from Preferred (LfP), Learning from Non-preferred (LfN), and Learning from Preferred and Non-preferred (LfP+N) settings on the DSRL benchmark. Across these varied settings, PSIL maintains a favorable balance between reward and safety trade-off. We support our empirical results with theoretical and empirical analysis of the trajectory-weighting mechanism.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.