EmboDirector: Directing 4D Human–Scene Interactions from a Single Image
Abstract
Full-body human–scene interaction data with consistent geometry, motion, and contact is scarce and costly to capture. Video models acquire rich interaction priors through large-scale pretraining, yet generated interactions can exhibit object-orientation drift that direct reconstruction would reproduce. We present EmboDirector, which reconstructs structured 4D interactions from videos generated using a single indoor image and an interaction description. A multimodal large language model (MLLM) Director corrects the initial recovered layout, assigns frame reliability to reduce the influence of video drift, and organizes estimated contact candidates into a spatiotemporal interaction graph. The graph specifies interaction relationships and their active intervals, guiding joint optimization of human and object trajectories alongside reliable visual evidence. The input image anchors layout corrections throughout. Comparisons with HOI generation and video-based reconstruction baselines show improved contact rate, mean penetration, and object alignment; an ablation examines graph-conditioned optimization. The reconstructed scenes support physics simulation and geometry-guided video synthesis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.