Any2Ego: Egocentric Viewpoint Translation Taking Textual Modality as Anchor
Abstract
Egocentric video is the format in which embodied agents see the world and is difficult to collect, whereas third-person recordings of the same activities are abundant. Translating these recordings requires recovering both the geometry and the appearance of interactions from the first-person viewpoint. Spatial priors constrain geometry, but leave appearance and scene details underspecified; using a text prompt alone does not determine which details supplement these priors or where they should influence generation. We introduce Any2Ego, which learns a first-person textual anchor complementary to hand–object interaction (HOI) masks and controls its influence through HOI-conditioned spatial routing. A Qwen3-VL-8B converter learns structured descriptions through supervised caption distillation with sentence-level weighting, followed by generation-grounded reinforcement that rewards agreement between added sentences and rendered previews. A lightweight gate then uses HOI features to modulate caption cross-attention at each video token. On held-out Ego-Exo4D clips, Any2Ego substantially improves perceptual fidelity over our reproduced EgoExo-Gen, better preserves viewpoint-dependent scene structure, and produces interactions whose appearance and temporal evolution more closely match the target egocentric videos.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.