acceptodds
Under review as a conference paper at ICLR 2027

CRAVE: Recovering Progress from Repeated Demonstrations for Contact-Rich Robot Post-Training

Abstract

Progress-conditioned robot post-training can emphasize the brief transitions that determine contact-rich task completion, but its supervision typically requires manual semantic stages, a separately trained value estimator, or elapsed-time proxies that confuse time with task advancement. We introduce Cross-Episode Recurrent Abstraction via Multimodal Evidence (CRAVE), an offline relabeler that recovers an episode-consistent progress coordinate from visual–state configurations that recur across demonstrations. Unlike framewise temporal proxies, CRAVE couples assignments over each complete episode and converts future progress into ordinal training conditions for π0.5 while leaving its objective and deployed architecture unchanged. Mechanistically, on 300 duration-balanced held-out garment-folding episodes, trajectory coupling localized human-referenced phase boundaries 12.72 s more accurately than framewise assignment (95% paired-bootstrap confidence interval, 11.38–14.09 s). Downstream, across 180 real-robot rollouts from nine matched policy checkpoints, Direct CRAVE completed 11/20, 15/20 and 16/20 trials at 150, 300 and 450 demonstrations, versus 10/20, 14/20 and 16/20 for the strongest matched comparator (plain behavior cloning or an adapted historical χ0 Stage Advantage labeler). Across tasks, the same relabeling interface transferred to nail painting and low-data writing: observed mean four-stage completion was 35% for CRAVE versus 30% for SFT over 20 nail-painting episodes per route, and the CRAVE-labeled writing policy completed 16/19 unique rollouts after training on 55 demonstrations. Operationally, a common eight-A100 engineering projection placed CRAVE label construction at 3.7–5.6 min versus 66–137 min for inference and export from an existing Stage Advantage checkpoint. Repeated demonstrations can therefore provide progress supervision for contact-rich VLA post-training without human phase labels, on-policy collection, or a separately trained reward, value or advantage estimator.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.