acceptodds
Under review as a conference paper at ICLR 2027

Success Rates Are Not Model Properties: Controlled Diagnostics for VLA Evaluation

Abstract

Success rates on standard manipulation benchmarks are the primary evidence for comparing vision-language-action (VLA) models. But a success rate is not a model property: it is measured under conditions fixed by the benchmark or by the configuration distributed with the model. We show that three of these conditions, changed with the model fixed, move success by amounts that are large for some models and small for others. The three are (1) which initial states are tested, (2) what instruction the model receives, and (3) how many predicted actions it executes between observations, the execution horizon. We test multiple VLA models on LIBERO and LIBERO-Plus, and execution also in SimplerEnv. (1) For initial states, X-VLA and π0.5 both succeed on about 98% of LIBERO-Object's standard initial states at the same execution horizon, but on 22.1% and 89.5% of placements outside these states on which a separate VLA model succeeds at a high rate. After task-specific fine-tuning of π0.5, success on the same placements falls by more than the standard score changes, for every recipe we tested. (2) For the instruction, removing the parameters that LIBERO-Plus appends to most instructions raises the success of the two SmolVLA models by up to 51.7 percentage points and changes the other eight models by at most 8.2 points. For these two models, the delivered text, not just the scene, drives success. Appending two ordinary words to standard instructions also lowers their success, even though both models perform another task when given that task's instruction. (3) For execution, a horizon of ten instead of each model's default raises the standard score of five of the six models on LIBERO and also changes success on the placements from (1), while in SimplerEnv each tested configuration does worst at its shortest horizon. Because models respond differently, these conditions move the margins between models and sometimes reverse which model scores higher, so a comparison made under one setting does not necessarily hold under another. We give an evaluation procedure that reports each model's response to these conditions and states when a comparison holds, asking not only how often a model succeeds, but under which conditions it does.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.