acceptodds
Under review as a conference paper at ICLR 2027

OmniHOI: Generalizable Human-Object Interaction Generation through Explicit Interaction Modeling

Abstract

Synthesizing human-object interaction (HOI) is valuable for applications in robot imitation learning and AR/VR animation. However, existing methods are ineffective in handling multiple-object interactions, or otherwise specialized to particular object configurations. Meanwhile, they fail to generalize to object categories or trajectories unseen during training. We present OmniHOI, a single HOI framework that jointly tackles single-object, articulated-object and multiple-object interactions. It devises a hierarchical generation pipeline, consecutively outputting object motion from text instructions, hand motion from object motion, and full-body motion from both object and hand motion. To reinforce interaction information, OmniHOI introduces interaction-aware self-conditioning and interaction-informed loss for diffusion models. OmniHOI also introduces multi-object test-time optimization to enhance hand contacts with multiple objects. OmniHOI sets a new state of the art on GRAB, ARCTIC and HIMO dataset, achieving stronger motion quality, text alignment, and hand–object interaction accuracy. Conditioned on unseen object categories or object trajectories, OmniHOI demonstrates stronger out-of-distribution generalization than prior work. Figure 1 summarizes our findings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.