Teaching Vision-Language Models to Evaluate Actions Leads to Better Manipulation Policies
Abstract
Vision-language-action models rely on pre-trained vision-language backbones, yet further pre-training with robot-specific descriptive objectives yields little downstream benefit. We therefore investigate whether an evaluative objective that distinguishes viable actions from near-misses and failures transfers more effectively to control. As a first instantiation, we consider object grasping and introduce Action Assessment Pre-Training (AAPT), which trains the backbone to autoregressively generate a physical reasoning trace followed by a scalar viability score for each candidate grasp. After pre-training, the assessment interface is discarded and only the backbone initializes the downstream policy. We separate the objective from both its data and its optimizer. On identical candidates, a predictive control supervised by simulated collision and stability outcomes already improves downstream control. Replacing its binary targets with graded evaluation under matched optimization improves further, and ranking refinement yields on zero-shot LIBERO-Plus and on CALVIN ABCD. Permuting complete trace-and-score targets across candidates causes performance to fall below no pre-training, showing that candidate-quality alignment matters. AAPT's transfer advantage also grows from to points as demonstrations become scarce. Discarded alternatives are therefore useful supervision, and their graded relative ordering is what transfers to manipulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.