GainPilot: Steering Vision-Language-Action Policies with Gain-Triggered Replanning and Bounded Action Correction
Abstract
Action-chunked vision-language-action (VLA) policies can accumulate execution errors, but detecting deviations alone does not determine whether replanning helps. We introduce GainPilot, which combines gain-triggered replanning with bounded post-sampling correction of a base policy. A dynamics-informed trigger predicts the terminal-success gain of replacing the active queue with a subsequently sampled and corrected plan. An anchored critic jointly learns absolute success, anchor-relative gains, and outcome differences between paired local perturbations. One action-gradient computation then supplies a coordinate-aware correction direction, bounded separately for translation and gripper commands. Across 13 backbone–horizon configurations on Meta-World MT50, spanning , SmolVLA, and X-VLA, GainPilot improves success by 10.6–26.0 percentage points over the base policies. For at , success rises from 61.6% to 74.2%, compared with 69.4% for gain-triggered replanning alone. On LIBERO with 10-shot fine-tuning, mean success rises from 79.00% to 90.75%; across nine real-robot tasks, it increases from 53.3% to 73.9%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.