acceptodds
Under review as a conference paper at ICLR 2027

Language-Critique Imitation Learning from Suboptimal Demonstrations

Abstract

Prior work on imitation learning from suboptimal demonstrations typically extracts supervision as a scalar per state or trajectory (e.g., a confidence estimate, discriminator score, or importance weight), which ranks behavior but cannot express task stage, what went wrong, or how to correct it. We propose language-critique imitation learning, an offline algorithm that learns policies with language labels based on expert and suboptimal demonstrations. Our method constructs language labels describing task progress, action optimality, and movement guidance, produced either by a lightweight feature-based generator or an off-the-shelf VLM. We then introduce a language-critique loss to train the policy using language signals without compressing them to scalars. We instantiate this loss for behavior cloning and diffusion policies, yielding LC-BC and LC-DP. We theoretically show that the proposed objective upper-bounds the expert performance gap under standard assumptions. Across seven continuous-control tasks in navigation, driving, and manipulation, LC-BC and LC-DP achieve the highest average success rates in their policy families, improving average success by 8.2% (BC) and 4.3% (DP). Further experiments demonstrate robustness to noisy expert selection and, on one task, comparable performance to full-data BC and diffusion policies using only 10% and 5% of the expert data, respectively. These results show that language critiques provide a practical and reliable learning signal for suboptimal demonstrations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.