acceptodds
Under review as a conference paper at ICLR 2027

ReCycle: Training Visuomotor Diffusion Policies on Unlabeled data with Cycle-Consistency

Abstract

Visuomotor diffusion policies are trained by imitation on action-labeled demonstrations, which are orders of magnitude scarcer than action-free video. Classical control couples forward and inverse dynamics through a shared notion of action: the action that explains a transition must also produce it. Many recent diffusion policies already represent this structure: their policy, forward-dynamics, and inverse-dynamics roles are conditionals of one learned distribution, linked by a shared latent action. Yet imitation constrains these conditionals only on the demonstrations, so the denoised trajectory space is optimized for imitation accuracy alone. We constrain it to also encode action-consistent transition structure by enforcing that latent actions inferred from a transition are sufficient for forward evolution under the same policy latent space. The resulting cycle-consistency objective turns action-free video into supervision for the policy without fabricating action labels. Enforced naively on out-of-domain video, however, it degrades performance, for two reasons: unlabeled transitions may not be realizable by the robot, and the cycle gradients conflict with the supervised signal. We address the first with slot-wise latent mixup, which brings unlabeled transitions in-domain while leaving the action variable untouched, and the second with a staged schedule and batch-relative variance reweighting. The resulting recipe, ReCycle, plugs into Cosmos Policy (CP), Unified Video Action Model (UVA), and Video Policy (VP) without changing inference. Using action-free video from physics simulation, other robots, or humans, ReCycle improves CP, UVA, and VP by up to , , and  pp on LIBERO-10 and RoboCasa, and improves CP's real-robot success by  pp on average. Its gains also persist under shifted friction, contact softness, and damping. Gains are largest when the unlabeled transitions contain dynamics relevant to the target task and shrink or reverse when the transition statistics are mismatched.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.