CompAct: Compositional 3D Human-Object Interaction Synthesis
Abstract
We rarely interact with a single object; we carry, hold, and manipulate several at once. Most methods on 3D interaction synthesis ignore multi-object interactions, as this needs new data capture or augmentation of many (ideally all) possible combinations; however, this does not scale. Our insight is that multi-object interactions are compositions of single-object ones. Thus, we represent interactions at the body-part level through interaction fields (InterFields) that encode contact and proximity. We exploit this in CompAct, a novel Flow-Matching model that, given a text prompt and body and object shape, jointly generates body pose, object pose, as well as body-part and object InterFields. At test time, CompAct predicts proxy single-object interactions and dynamically composes a multi-object one, using classifier-free guidance. To this end, composition exploits blending weights that naturally arise from the per-part InterFields of proxy interactions. Note that this is not a naive linear combination; instead, it is informed by a body prior arising from the classifier-free guidance, so that the composed interaction looks natural. Moreover, note that CompAct uses decoupled timesteps per input part; this supports interaction editing as a downstream application. CompAct is the first model to synthesize multi-object interactions by training only on single-object ones. Experiments on seven single- and multi-object benchmarks show that CompAct outperforms the state of the art in interaction fidelity, and produces multi-object interactions that are perceived as natural looking. Code and models will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.