acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical Vision Language Action Model Using Success and Failure Demonstrations

Abstract

Vision–Language–Action (VLA) models are typically post-trained on teleoper- ated successful demonstrations, while the many failed attempts that occur natu- rally during data collection are discarded. Those failures record where and how a policy is fragile, and we show they become a usable learning signal once plan evaluation is posed as a reach–avoid problem. We introduce VINE, a hierarchi- cal VLA that separates high-level reasoning (System 2) from low-level control (System 1) under a semi-Markov option formulation. Its central component is a subgoal-level value function that we prove equals the first-exit probability of reaching the goal set before the failure set. Because that quantity is defined by outcomes rather than by expert actions, it can be regressed directly from mixed- quality offline data, with no reward shaping and no online interaction. System 2 uses this value to run a batched tree search over a text-serialized scene-graph abstraction, proposing subgoal transitions, predicting their successor states, and pruning brittle branches before any motion is executed. System 1 then executes the selected subgoal sequence without altering the agent’s motor skills. Across simulated manipulation, the SIMPLER benchmark, and a real 6-DoF arm, VINE improves success most sharply under distribution shift (1.6→over the ω0 baseline on unseen plug insertion and 5.5→on unseen long-horizon real-world packing), suggesting that failure data discarded during collection can improve execution un- der shift.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.