acceptodds
Under review as a conference paper at ICLR 2027

Steering Generative Robot Policies with Optimal Transport Maps

Abstract

Generative policies have demonstrated significant potential in robotic manipulation. Despite acquiring a rich repertoire of motor skills, they remain brittle under train-test shifts: even mild changes in geometry or context can cause execution failures, limiting scalable deployment in the real world. While retraining or fine-tuning offers a remedy, it often incurs impractical costs and may impair the pretrained generalist capabilities. To this end, we propose LOTUS, a training-free framework based on Optimal Transport (OT) that adapts the behavior of frozen policies during inference. Following the VLM-guided paradigm, LOTUS first synthesizes stage-wise reward functions to evaluate candidate actions. Unlike previous efforts that use rewards for best-of-N selection or gradient perturbation, we construct virtual trajectories along high-reward directions, which can be viewed as samples from a task-aligned distribution. Our key idea is to project the original policy distribution toward such reward-induced regions while preserving the learned action prior. This can be formulated as an OT problem with motion discrepancy as the transport cost, and we further relax the target marginal to avoid over-steering toward implausible actions. We evaluate LOTUS on both simulated and real-world manipulation tasks, showing substantial improvement against diverse perturbations. The code is provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.