acceptodds
Under review as a conference paper at ICLR 2027

WorldZoo: Compositional World Editing for Generalizable Robot Learning

Abstract

Imitation learning for robotic manipulation is limited by the high cost of collecting demonstrations and insufficient data diversity. By reusing existing demonstrations, data augmentation offers an effective way to broaden training data coverage under a limited collection budget. However, composing multiple edits without disrupting task-relevant interactions and selecting augmentations that benefit a target task remain challenging. We present **WorldZoo**, a unified world-editing framework that leverages robot demonstrations and human egocentric videos to generate augmented data through composable edits along three dimensions: embodiment, interaction, and scene. WorldZoo coordinates embodiment retargeting, visual editing, and trajectory adaptation to maintain consistency among observations, actions, and task-relevant contact relationships across diverse robotic arms and dexterous hands. Object editing adapts the demonstrated contacts to new object geometry within a range compatible with the original manipulation, expanding object diversity while preserving the demonstrated task. UniAug composes these edits within each sample as a reusable default recipe. We extend existing robotic simulation benchmarks to evaluate generalization across all three dimensions, systematically assess the effectiveness of augmentation along each dimension, and evaluate the training value of augmented data by fine-tuning generalist robot policies. In real-world experiments, UniAug matches or exceeds applying edits separately with one-sixth of the samples.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.