acceptodds
Under review as a conference paper at ICLR 2027

HARP: Hindsight Reflection Autonomously Improves Robot Policies

Abstract

Real-world robot policies must operate in unseen settings and perform tasks beyond the support of their pretraining data. While scaling up pretraining has enabled broadly capable robot policies, learning autonomously from experience allows policies to continually improve as they encounter new scenarios. A central challenge to achieving this is incorporating prior knowledge in the learning process, allowing robots to interpret their mistakes and determine how to improve their behavior, as humans do. To address this, we introduce HARP, or Hindsight-guided Adaptation for Robot Policies, a general framework for autonomous improvement of language-conditioned robot policies through hindsight reflection. HARP uses the prior knowledge of VLMs to reflect on robot experience, determining which decisions led to successful outcomes, caused failures, or introduced risks even when no failure occurred. This feedback is converted into subtask-level corrections, aligned with an overall task strategy. These corrections supervise ongoing updates to a high-level planner, improving its ability to guide the language-conditioned policy through subsequent interactions. We instantiate HARP for autonomous driving and demonstrate improvements over base policies on challenging routes from Bench2Drive and Fail2Drive.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.