From Evaluation to Recovery: Value-Guided Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models learn to generate actions from demonstrations, but the imitation objective does not teach them to compare those actions. Executing a poor sample can therefore leave the robot in an unfamiliar state from which success is difficult. We introduce VaSeR, which equips an existing VLA with an action-conditioned value for selection and recovery. A value token in the policy’s own backbone scores each candidate through the vocabulary head in one forward pass. It learns from demonstrations alone, without failure data: each demonstrated action is paired with a constructed negative whose pseudo-label decreases with its deviation from the demonstration. At inference, VaSeR executes the highest-valued of several sampled candidates and records its value before execution. When these values drop or stagnate, the robot physically retreats through its recent poses and resumes. On ten RLBench tasks, VaSeR raises success from 60.2% to 76.5% on π₀.₅ and from 52.3% to 65.3% on HybridVLA; on two real-robot tasks, the corresponding improvements are from 58.3% to 73.3% and from 53.3% to 78.3%. Held-out evaluations show that constructed negatives improve both candidate ranking and episode-level alarm accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.