acceptodds
Under review as a conference paper at ICLR 2027

Learning from Last-Mile Failures in Fine-Grained Robotic Manipulation

Abstract

Fine-grained robotic manipulation operates at tight tolerances; therefore, real-robot data collection inevitably yields many failed trials. Standard practice discards these failures or compresses each into a scalar reward, preference, or advantage. We take a global view: all trials form one record of how a task is performed, and success is a label set by the task tolerance. This record reveals a sharp regularity: most failures are **last-mile**, reproducing the successful motion and departing only where contact decides the outcome. We show that this structure turns a trajectory-level outcome label into a precise correction signal. (i) Every two-weight guidance rule over success-, failure-, and null-conditioned predictions, including existing attraction, repulsion, and two-branch rules, acts along one contrast direction and differs only in its gain. (ii) This direction aligns with the task-critical correction whenever a measurable scaffold ratio is below one, as for last-mile failures, so one label per trajectory localizes the error without step annotations. On this basis, we propose **CALM** (**C**ontrast **A**t the **L**ast **M**ile), which trains a world action model with outcome tokens and applies contrastive denoising guidance that attracts actions toward success and repels them from failure. On six real-robot tasks, CALM improves average success by 12.8 points over the standard baseline, with gains on every task, and loses this gain once off-scaffold failures raise the scaffold ratio above one, as the validity condition predicts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.