acceptodds
Under review as a conference paper at ICLR 2027

All-in-One Unified Egocentric World Simulator

Abstract

Egocentric and exocentric videos are two views of the same physical event: the first-person view captures intent-driven hand–object interaction, while the third-person view reveals the camera wearer's full-body motion and the surrounding scene. Most cross-view generators translate between the two in one predetermined direction, so different directions and temporal completions require separate models. In this paper, we introduce EgoU, a unified generative framework that places one egocentric and up to three exocentric views on a shared video canvas and expresses diverse generation tasks through conditioning masks and optional text prompts. Fine-tuned from a pretrained video diffusion transformer (DiT), EgoU supports exocentric-to-egocentric generation with or without an egocentric seed frame, egocentric-to-exocentric generation, temporal extension, temporal interpolation, text-driven egocentric synthesis, and joint cross-view generation within a single set of parameters. We incorporate view-type embeddings to differentiate perspectives, condition on scene geometry and the wearer's head trajectory through a dedicated cross-attention branch, and apply an auxiliary pose loss to supervise exocentric human motion. To train and evaluate this formulation, we build EgoExoVideo-600K, a corpus of 604,113 synchronized ego–exo clip pairs, and EgoU-Benchmark, which scores all tasks on the same held-out clips. To our knowledge, EgoU is the first single model to support all of these tasks through a shared interface, and it matches or surpasses task-specific baselines. EgoU is open-sourced anonymously at huggingface.co/EgoVideo/EgoU-5B.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.