Closed-Loop Adaptation of Frozen Vision–Language–Action Models
Abstract
Execution errors can prevent vision-language-action (VLA) policies from completing tasks they solve under nominal robot dynamics. Correcting these errors changes both motion and the observations driving subsequent policy commands. We present an execution interface, calibrated from fault-free motion, that estimates one additive offset per modeled command coordinate and applies bounded correction with frozen policy weights. To analyze the resulting closed-loop system, we introduce crossed command-stream replay, which separates direct correction from changed-command effects while retaining observer feedback. Geometric and finite-horizon analyses relate physical repair to response direction, correction authority, motion-error persistence, and information delay. Our simulation study spans , , OpenVLA-OFT, and GR00T N1.5/N1.7 in four robot interfaces. Across four LIBERO benchmark suites, success under additive command faults rises from 58/220 to 121/220 over two sampler seeds on the same 110 task/state cases held out from calibration and method selection; healthy success changes from 209/220 to 213/220. With four previously untested rotation-offset vectors in forty new task/state cases, adaptation increases success from 48/160 to 125/160; healthy success is 37/40 with and without adaptation. The averaging gain selected for development is 0.16 versus 0.08 in the initial-state holdout. Panda-arm calibration transfers to OpenVLA-OFT and GR00T N1.7. With commands fixed, correction reduces median endpoint deviation from 13.7 to 7.8 mm. Replay shows that changed commands can reduce or increase reference-relative translation path-deviation cost. Separate predictor interventions distinguish identification accuracy from task recovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.