acceptodds
Under review as a conference paper at ICLR 2027

Free-View Manipulation via Dual-Flow Generation Policy: Canonical Images and Actions

Abstract

Visuomotor policies have become the mainstream action generation head in robot manipulation research. However, in open-world environments, the camera viewpoint varies arbitrarily and breaks the perception–action spatial alignment, leading to poor generalization and unstable execution. To address this issue, we propose a novel Dual-Flow Generation Policy (DFGP) that implicitly learns viewpoint-invariant visual representations to adapt to free-view observations and guide accurate manipulation action generation. Our framework consists of two collaborative pathways: an image mapping pathway that infers canonical-view images from free-view inputs via flow matching, and an action generation pathway that predicts robot actions conditioned on the shared canonical visual features. We further introduce a Bi‑directional Cross‑Attention (BiXT) module to adaptively fuse multi-source features for stronger cross-view action generalization. Extensive experiments conducted on two simulation benchmarks (Free-View and LIBERO) —as well as on physical world robotic platforms demonstrate that DFGP achieves statistically significant improvements over prior visuomotor policies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.