PriorHOSI: Composing Motion Priors for Human–Object–Scene Interaction Generation
Abstract
Human–object–scene interaction (HOSI) generation aims to synthesize coordinated human and object motion compatible with the surrounding scene. Existing HOSI datasets remain scarce, while high-quality, diverse human–object interaction (HOI) and human–scene interaction (HSI) datasets offer complementary supervision. The challenge is to adapt human–object motion to scene geometry while preserving the contact relationships that coordinate the interaction. We propose PriorHOSI, a framework that learns a human–object interaction prior (HOIPrior) and a human–scene interaction prior (HSIPrior) separately and composes them at inference time. The key idea is to use HSIPrior’s scene-adapted human motion to guide joint updates of the body and object through HOIPrior. Diffusion noise optimization realizes these updates by fitting the generated motion to the scene-adapted target while constraining hand positions in the object’s coordinate frame. To support repeated updates, HOIPrior uses relation-aware consistency distillation to reduce sampling steps and retain hand–object geometry. To better exploit scene geometry, HSIPrior uses body-part-aware residual refinement to propose scene-aware motion adjustments during editing. Experiments on the InfBaGel benchmark demonstrate that PriorHOSI achieves strong task performance and hand–object contact while improving human–scene compatibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.