acceptodds
Under review as a conference paper at ICLR 2027

VLA-Retain: Benchmarking Cross-Dimension Effects of Vision Tuning in VLAs

Abstract

Vision-language-action (VLA) models transfer broad visual and semantic knowledge from large-scale pretraining into robot control, but their manipulation performance remains sensitive to changes in observation conditions, which recent work addresses by fine-tuning the visual module on target-domain data. However, it remains unclear how such adaptation affects robustness beyond the target condition, as existing robustness benchmarks only evaluate models without fine-tuning and adaptation studies only report the target gain. Measuring this cross-dimension effect requires separating target gains from off-target changes, isolating adaptation effects from architectural differences, ruling out generic forgetting, and comparing tuning choices. Therefore, we introduce **VLA-Retain**, a paired benchmark that compares five open VLA models, tuned on Camera-view or Sensor-noise data, with their own base checkpoints on the same 10,030 LIBERO-Plus episodes spanning seven perturbation dimensions under varied training data and tuning strategies. Our study reveals four findings. First, targeted vision tuning is not local, and its target gain does not predict retention. Mean off-target gain is negative in 8 of 10 model–domain configurations. Under Camera-view tuning, X-VLA gains percentage points on the target but loses points on average elsewhere, whereas OpenVLA-OFT gains and points, respectively.Second, these regressions are not explained only by performance loss on the same tasks without perturbations. Third, broader training coverage raises overall gain over target-specific training but does not eliminate dimension-level regressions. Fourth, tuning strategies are model-dependent, as DELTA-style feature preservation raises X-VLA's mean off-target gain to points but lowers OpenVLA-OFT's to . Overall, these findings show that target-domain success alone is insufficient and highlight cross-dimensional retention as a key criterion for targeted VLA adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.