IMAGIN-4D: Image-Guided Controllable Interaction Generation
Abstract
Generating human-object interactions (HOI) is central to character animation, robotics, AR/VR, and embodied AI. Existing methods use text, object geometry, and sparse waypoints to control actions and object trajectories, but leave grasps, approach directions, body and object poses, contacts, and relative layouts ambiguous. We use a reference image to specify the desired interaction snapshot. However, a single global image representation conflates distinct cues and uses the same visual evidence to condition every frame. We introduce IMAGIN-4D, a diffusion-based HOI generator with spatio-temporal image conditioning. For spatial conditioning, supervised interaction-state tokens encode body pose, object pose, body-object contact, and spatial relationships at the depicted frame. For temporal conditioning, frame-aware tokens query image patches per generated frame to retrieve visual cues. Role-aware conditioning balances these signals: text, waypoints, and interaction-state tokens use separate AdaLN streams, while motion tokens cross-attend to frame-aware visual tokens. Since HOI motion datasets lack paired images, we render synthetic references from FullBodyManipulation (FBM) and introduce an image-adherence metric measuring agreement with the reference snapshot. Experiments on FBM and BEHAVE show that IMAGIN-4D improves fine-grained interaction control over single-token and uniformly image-conditioned baselines while preserving waypoint-following and motion quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.