acceptodds
Under review as a conference paper at ICLR 2027

HOHI4D: Scalable Synthesis of 3D Human-Object-Human Interaction Data

Abstract

3D human-object-human interaction (HOHI) modeling remains underexplored due to the lack of scalable datasets and the difficulty of acquiring accurate multi-agent interaction data. Existing approaches rely on controlled capture setups or strong assumptions on object pose and contact, limiting their applicability in complex scenarios. In this work, we propose a fully automated and controllable pipeline for large-scale HOHI data synthesis, combining prompt-driven video generation with SAM3D-based reconstruction. Our framework enables diverse interaction generation, reduces object drift via geometry-aware refinement, and removes reliance on predefined object templates. Based on this pipeline, we construct HOHI4D, a large-scale dataset with over 5,000 sequences and 13 hours of motion. Furthermore, we introduce a structured generative model for HOHI. We design HOHI-VAE to learn disentangled latent representations of individual motion, object dynamics, and global interaction, and build HOHI-Diff, a latent diffusion model operating in this structured space. Our approach enables stable and coherent generation of complex multi-agent interactions, providing a scalable benchmark for HOHI modeling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.