Vision-Grounded Energy Critics for Physically Feasible Vision-and-Language Action Policies
Abstract
We introduce Vision-Dominated Energy Learning (VEL), a lightweight, backbone-agnostic framework that augments any frozen or fine-tuned vision-and-language action (VLA) policy with a compact energy head. The energy head maps the hidden representations of the VLA and a candidate action to a scalar energy value, where lower energy indicates higher compatibility between the action and the current scene. Training uses a contrastive objective on demonstration data, and inference applies a single gradient-based residual correction that shifts the proposed action toward a lower-energy alternative. This procedure requires no environment rollouts, no additional data collection, and no modification of the VLA backbone. At the population optimum, the contrastive objective recovers the log density ratio between the conditional expert action distribution and the action marginal, and a single gradient step is guaranteed to reduce the learned energy under local smoothness conditions. Over five training seeds of the energy head, VEL raises the LIBERO-Long success rate of OpenVLA-OFT from 94.4% to 95.8% and that of from 92.4% to 93.6%, which removes 24% and 16% of the residual failures, while the shorter suites show smaller changes of mixed sign. Controlled comparisons show that search over the same energy, a direct correction head of matched capacity, and a score-matching energy head do not reproduce the long-horizon gain. On a physical screw-insertion task with a backbone, VEL raises full-task success from 35% to 50%, and the gain concentrates in the precision-demanding insertion stage. The energy head introduces fewer than 4.0M additional parameters and 4.61 GFLOPs per correction step, an increase of 0.4% over OpenVLA-OFT and 5.2% over .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.