Language-Critique Imitation Learning from Suboptimal Demonstrations
Abstract
Prior work on imitation learning from suboptimal demonstrations typically extracts supervision as a scalar per state or trajectory (e.g., a confidence estimate, discriminator score, or importance weight), which ranks behavior but cannot express task stage, what went wrong, or how to correct it. We propose language-critique imitation learning, an offline algorithm that learns policies with language labels based on expert and suboptimal demonstrations. Our method constructs language labels describing task progress, action optimality, and movement guidance, produced either by a lightweight feature-based generator or an off-the-shelf VLM. We then introduce a language-critique loss to train the policy using language signals without compressing them to scalars. We instantiate this loss for behavior cloning and diffusion policies, yielding LC-BC and LC-DP. We theoretically show that the proposed objective upper-bounds the expert performance gap under standard assumptions. Across seven continuous-control tasks in navigation, driving, and manipulation, LC-BC and LC-DP achieve the highest average success rates in their policy families, improving average success by 8.2% (BC) and 4.3% (DP). Further experiments demonstrate robustness to noisy expert selection and, on one task, comparable performance to full-data BC and diffusion policies using only 10% and 5% of the expert data, respectively. These results show that language critiques provide a practical and reliable learning signal for suboptimal demonstrations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.