A Training-Free Approach for Adapting Flow-Matching Policies to Spatial Shifts
Abstract
A deployed pre-trained policy often has to operate in a fixed spatial layout that differs from the one it was trained on. This spatial shift places the policy outside its training distribution, where it frequently fails. We seek a training-free adaptation that raises the success rate at run time using only the policy's own attempts at that layout: no demonstrations, no weight updates, only the outcomes of its own attempts. A success recorded there represents a complete trajectory that solves the layout; with the weights frozen, that trajectory can only be exploited through the sampler. For a chunked flow-matching policy, the sampler offers three entry points: the measure it starts from, the time at which it enters the flow, and the action it emits. We inject the stored success at all three points. Specifically, warm-started denoising re-enters the flow from the success under the current observation; a density gate, using the density defined by the same velocity field, returns control to the policy once the re-solved chunk is no longer one the policy would produce; and an action-space residual search adds a constant offset fitted using terminal outcomes alone. We consolidate these into WDRS, a plug-and-play adaptation framework, and conduct extensive empirical validation. On LIBERO, WDRS raises single-draw success from to at the smallest of three jitter scales and matches or exceeds a per-site weight-update baseline in five of six suite-by-jitter cells; it carries over to two further frozen flow policies (up to ), and to the Layout perturbations of LIBERO-plus () and RoboLab (, warm-started only). This demonstrates that WDRS is a highly transferable and readily deployable solution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.