acceptodds
Under review as a conference paper at ICLR 2027

Teaching Vision-Language Models to Evaluate Actions Leads to Better Manipulation Policies

Abstract

Vision-language-action models rely on pre-trained vision-language backbones, yet further pre-training with robot-specific descriptive objectives yields little downstream benefit. We therefore investigate whether an evaluative objective that distinguishes viable actions from near-misses and failures transfers more effectively to control. As a first instantiation, we consider object grasping and introduce Action Assessment Pre-Training (AAPT), which trains the backbone to autoregressively generate a physical reasoning trace followed by a scalar viability score for each candidate grasp. After pre-training, the assessment interface is discarded and only the backbone initializes the downstream policy. We separate the objective from both its data and its optimizer. On identical candidates, a predictive control supervised by simulated collision and stability outcomes already improves downstream control. Replacing its binary targets with graded evaluation under matched optimization improves further, and ranking refinement yields on zero-shot LIBERO-Plus and on CALVIN ABCD. Permuting complete trace-and-score targets across candidates causes performance to fall below no pre-training, showing that candidate-quality alignment matters. AAPT's transfer advantage also grows from to points as demonstrations become scarce. Discarded alternatives are therefore useful supervision, and their graded relative ordering is what transfers to manipulation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.