EgoTACO: Egocentric Task Abstraction for CrossEmbodiment Object Manipulation in Editable Worlds
Abstract
Egocentric video records the goal, contact roles, event order, and object-state changes of everyday manipulation. Direct video-to-robot retargeting binds this evidence to a human execution that can become invalid when robot morphology or object physics changes. We introduce EgoTACO, a task-conditioned video generation framework that separates a source into completed background, human-operation, and object layers, infers an editable task latent from the latter two, and regenerates the full scene under target-object, physics, and robot-embodiment conditions. The latent specifies what must change in the world without prescribing a human wrist path. A simulation harness recompiles edited tasks and verifies kinematics, contact roles, and terminal state. Joint geometry and inertia adaptation eliminates the corresponding violations in 120 constructed interventions, while task-contact replanning completes all 60 cross-embodiment simulator goals. A conditional video predictor gains 3.79 dB motion-region PSNR over an equal-capacity unconditioned model. Missing end-to-end and policy results are explicitly marked in the tables.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.