acceptodds
Under review as a conference paper at ICLR 2027

Reference-Conditioned Cross-Embodiment Video Generation without 3D Robot Assets

Abstract

Human and robot demonstration videos provide a growing source of manipulation experience, yet each observation couples the interaction with the demonstrator's body. Reusing this experience across embodiments requires the target robot's configuration and occlusions to adapt to its morphology and the surrounding scene. This work presents Ego2Any, a reference-conditioned video diffusion model that generates corresponding realizations of egocentric manipulation interactions across robot embodiments. It represents the demonstrated interaction separately from target morphology, with morphology-conditioned temporal transport to infer target-specific body layouts throughout the interaction. These layouts guide trajectory-conditioned video synthesis to preserve the source interaction and scene context while remaining aligned with the source actions. Ego2Any supports embodiment changes within a shared model without requiring target meshes, kinematic models, or a separately prepared target initial frame at inference. We further construct EgoXPair, an action-paired dataset containing 1.0M retained videos across 10 target embodiments. Evaluations show that the generated demonstrations improve downstream policy learning, supporting cross-embodiment video generation as reusable manipulation experience. Our code and dataset will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.