acceptodds
Under review as a conference paper at ICLR 2027

Stage-Verifiable Latent Chain-of-Thought with Local Repair for Long-Horizon Robotic Manipulation

Abstract

Long-horizon robotic manipulation requires intermediate reasoning to remain consistent with observed task progress. Continuous latent plans, however, often lack stage semantics that can be checked during execution. We present a vision-language-action framework that aligns latent segments with demonstration-derived object relations and subgoal completion, verifies their expected outcomes against observations, and repairs the affected suffix after execution deviations. The framework retains a verified prefix while its still-required conditions remain valid and conditions regeneration on the current scene and task history. All model calls, including verification and repair, are included in a task-level inference budget. The method achieves measured mean success rates of 98.0% on LIBERO overall and 96.8% on LIBERO-Long, and 58.0%/23.0% under clean/randomized conditions on 50 RoboTwin 2.0 tasks. On LIBERO-RECOVER, enabling repair increases macro-averaged recovery success from 22.8% to 26.0%. Relative to full replanning, local repair reduces cumulative inference time from 18.6 to 14.9 seconds per episode, a 19.9% reduction, with recovery success of 26.0% versus 26.4%. These results characterize the recovery–computation trade-off under the evaluated conditions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.