COHERE4D: Training-Free Metric Alignment for Coherent 4D Human–Scene Reconstruction
Abstract
We present COHERE4D, a training-free framework that geometrically aligns the outputs of independently pretrained human and scene foundation models to reconstruct a coherent, metric 4D world from monocular videos. Recent unified human–scene reconstruction methods largely learn metric scene scale and human alignment from training data, but generalize poorly to out-of-distribution scenes. Instead, COHERE4D performs test-time geometric alignment of human and scene predictions using dense pixel-aligned correspondences and RANSAC-based robust estimation. This alignment provides mutual geometric grounding: near-metric human meshes anchor the scale of scene geometry and camera translation, while scene geometry in turn refines human translation to improve global human trajectories. This test-time alignment requires no additional training and uses RANSAC-based robust estimation to suppress outliers caused by occlusion, background contamination, and erroneous depth. Across benchmarks, COHERE4D significantly outperforms prior methods in human–scene alignment, metric-scale scene and camera pose reconstruction, and global human motion estimation. Our training-free alignment method adds only about 4 ms/frame, while the overall inference reaches up to 10 FPS.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.