acceptodds
Under review as a conference paper at ICLR 2027

CAHI4D: Cross-Age Human Interaction 4D Reconstruction from Monocular Video

Abstract

Age-dependent morphology exacerbates scale–depth ambiguity in monocular multi-person reconstruction, while adult-centric priors and occlusion of interaction cues further hinder cross-age interaction recovery. We present Cross-Age Human Interaction 4D Reconstruction (CAHI4D), which uses a reliable reference and validated interactions for geometric anchoring before fitting age-dependent shape and refining interaction motion. A vision–language model (VLM) proposes support and contact hypotheses from visual context; geometric and temporal checks then determine which relations can constrain reconstruction. Depth- and Interaction-Guided Geometric Anchoring (DIGA) uses a reliable scene-supported reference to align the shared scene scale, and then places a geometrically underconstrained participant through validated interactions and corrects body scale from the induced projection change. With fixed placement and overall body scale, shape fitting adapts age-dependent proportions. On the fixed geometry, Kinematic-Topology-Guided Interaction Correction (KTIC) maps directed surface corrections to participant-specific kinematic chains and reconciles corrections sharing joint degrees of freedom; contact and temporal terms constrain the same motion fit. CAHI4D achieves the lowest errors among the compared methods on Hi4D, 3DPW, and CMU toddler. In particular, on CMU, Root MPJPE and Joint-PA MPJPE fall by 21.0% and 49.1%, respectively. Code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.