acceptodds
Under review as a conference paper at ICLR 2027

Training-Free Adaptation of Generalist Robot Policies via Distributional Policy Steering

Abstract

Pretrained vision-language-action (VLA) policies possess broad motor skills but can struggle to apply them under semantic and spatial distribution shifts. Vision-language models (VLMs) offer complementary reasoning capabilities for out-of-distribution (OOD) inference, given its strong capability on semantic understanding and commonsense knowledge. However, how to formalize and integrate VLM reasoning into visuomotor action sampling while leveraging pretrained motor skills remains an open challenge. We introduce Distributional Policy Steering (DPoS), a training-free, inference-time steering method that provides distributional alignment guidance from VLM-derived spatial information to a frozen VLA policy. DPoS constructs a shell-supported target distribution around a reasoned spatial estimate and jointly guides sampled action chunks toward it. An energy function is designed based on the squared maximum mean discrepancy measures between the sampled action candidates cloud and the target distribution. DPoS enables the repulsion among sampled candidates, aiming for a better coverage among the action space. The attraction to the target distribution aims for a reasoned spatial target alignment for the task success. Evaluated on LIBERO-PRO simulator and 9 real-world OOD challenging tasks, DPoS achieves the highest success rate among all leading baselines. It improves over the base model by 44.4 and 19.9 percentage points on real-world and LIBERO-PRO tasks respectively, demonstrating distributional guidance as an effective way to integrate reasoning prior into pretrained robot control without additional model training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.