Collaborative Human-Object Interaction Reconstruction from a Monocular Video
Abstract
Collaborative human-object interaction (HOI) reconstruction requires recovering how several people manipulate a shared object over time. Existing video-based HOI reconstruction methods fall short in this scenario, as they neither align multiple human-object pairs in a shared 4D space nor reason about how contacts among multiple humans and the shared object evolve over time. To address this, we present COMBO, a 4D reconstruction method for reconstructing collaborative HOIs from a monocular video. Our key insight is that, in collaborative scenarios, the humans and the object are tightly coupled through contacts that evolve over time, and these time-varying contacts provide the constraint needed for joint reconstruction. We design a two-stage approach: human-anchored initialization first establishes a shared 4D space, using human motions to calibrate depth, and initializes the object. Subsequently, temporal contact-guided optimization builds a proximity-based temporal contact graph from 2D and 3D cues, without a learned contact predictor. Unlike static or single-pair contact formulations, our graph captures temporally varying contacts across participants, guiding the joint alignment of all entities in a shared 4D space. Experiments on the CORE4D benchmark show lower reconstruction errors than competing video-based methods. Our results on in-the-wild and generated videos demonstrate the generalization capability of COMBO. Code will be publicly available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.