acceptodds
Under review as a conference paper at ICLR 2027

MutualEgo: Estimating 3D Motion of Interacting Humans from Dual Egocentric Views

Abstract

Recovering the 3D motion of two interacting camera wearers is challenging because each egocentric view only partially observes the partner while moving with its own wearer. We introduce MutualEgo, a pair-centric diffusion framework that jointly estimates both participants' motion in a shared reference frame. By treating the interacting pair as a single prediction unit, MutualEgo integrates complementary evidence from both moving views while explicitly modeling their evolving spatial relationship. Its Reference-Aware Motion Tokenizer (RAMT) grounds dense DINOv3 features in calibrated camera geometry and uses anchor-guided attention to aggregate structured global and local conditioning tokens. To support training and evaluation, we introduce MutualEgo-10K, comprising 10,105 interaction sequences with rendered dual-egocentric RGB, calibrated camera geometry, and paired 3D motion. On MutualEgo-10K, MutualEgo achieves reconstruction and forecasting MPJPEs of 0.081 m and 0.287 m, compared with 0.261 m and 0.445 m for UniEgoMotion retrained on the same dataset. Ablations evaluate dual-view pair modeling and geometry-aware tokenization, while further experiments study training-data scale and adaptation to Harmony4D.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.