Do Language and Visual Perturbations Compound? Interaction Effects and Backend Sensitivity in Compact Vision-Language-Action Model Evaluation
Abstract
Robustness evaluations of vision-language-action (VLA) models typically report per-perturbation success rates from a modest number of episodes, treating each episode's outcome as fixed and each perturbation as separable. Both assumptions bear directly on the validity of the resulting benchmark numbers. We test them in MiniVLA, a compact VLA with roughly 1B parameters, on LIBERO-90, using a pre-specified factorial design that crosses instruction and visual perturbations over 7 tasks (630 episodes). Individual episode outcomes are fragile: across 70 pairs of episodes that share the task, initial state, and instruction and run greedily on the same machine and backend, shifting only the random seed changes the outcome in 26% of pairs and the episode length in 63%, while aggregate success moves by two episodes. A preliminary comparison suggests that switching the inference backend at fp32 has a similarly sized effect. Within the factorial, visual perturbation dominates. Mild lighting and camera changes reduce success from 57% to 20%, while paraphrased instructions have little effect. Mismatched instructions drive success on the original task to zero, so this model does not ignore language, unlike recent reports for other VLAs. For paraphrase combined with mild visual perturbation, the only pairing with room above zero success, we find no detectable compounding, and we show how floor effects cap the super-additive loss such a design can detect. Sampled rollouts are no more diverse than a greedy policy perturbed with small Gaussian action noise. We recommend that small-sample VLA robustness evaluations report seeds and the inference backend, treat outcomes at a fixed initial state as stochastic by repeating episodes and reporting interval estimates, and test perturbations jointly at strengths that leave headroom above zero success.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.