acceptodds
Under review as a conference paper at ICLR 2027

PSIL: Learning Safe Offline Policy from Preferred and Non-Preferred Demonstrations

Abstract

Offline imitation learning (IL) typically relies on expert (i.e., preferred) demonstrations and a large collection of supplementary trajectories for broader state-action coverage. However, existing approaches often struggle to effectively integrate the valuable signal from non-preferred demonstrations, which provide critical information about what the agent should avoid. We propose PSIL, Progressive Safe Imitation Learning, a self-training-based offline IL algorithm that can flexibly learn from preferred demonstrations, non-preferred demonstrations, or both. PSIL progressively propagates behavioral information from this small set of demonstrations to a large supplementary trajectory dataset by assigning preference-weights based on their similarity to preferred or non-preferred behaviors. These weights induce a relative ordering over trajectories and are used in weighted behavior cloning to emphasize preferred behavior while down-weighting non-preferred behavior. We evaluate PSIL in Learning from Preferred (LfP), Learning from Non-preferred (LfN), and Learning from Preferred and Non-preferred (LfP+N) settings on the DSRL benchmark. Across these varied settings, PSIL maintains a favorable balance between reward and safety trade-off. We support our empirical results with theoretical and empirical analysis of the trajectory-weighting mechanism.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.