4DHOISolver: Efficient and Accurate Human–Object Motion Reconstruction from Monocular Videos
Abstract
Internet videos capture a rich diversity of human activities, objects, and environments, offering a readily available source of interaction data for embodied intelligence. However, recovering accurate and physically plausible 4D human–object interactions (HOI) from these monocular videos remains challenging due to depth ambiguity, occlusion, and complex contact relationships. We introduce 4DHOISolver, a pipeline that combines automatic reconstruction with efficient manual correction. Our agent-based reconstruction method uses semantic interaction reasoning to guide geometric optimization, but challenging in-the-wild videos can still lead to reconstruction failures. These failures motivate an efficient and scalable manual repair method that corrects reconstructed motion through simple, sparse point annotations. We further propose a unified physical refinement method that corrects residual contact errors and interpenetration across diverse interactions without action-specific reward design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.