InterAssembly: Assembling Long-Horizon Human-Object Interactions via Retrieval, Optimization, and Generation
Abstract
Human-object interaction (HOI) datasets have advanced motion synthesis and embodied behavior modeling, but remain limited to short, isolated clips. This clip-level supervision cannot capture long-horizon activities that require temporally coherent manipulation of multiple objects and consistent scene evolution, essential for embodied intelligence. Collecting such data is prohibitively expensive, and no existing method generates long-horizon, multi-object HOI activities under today's data budgets. We introduce InterAssembly, an agentic system that assembles activity-oriented long-horizon HOI animations from an HOI database of short captured clips. Given only a 3D scene and this HOI database, an LLM agent iteratively proposes the next sub-action, retrieves candidate reference clips by structural scene similarity and selects among them by semantic fit to the proposed action, and adapts object placements to the target scene using collision-aware feedback. An interaction mesh retargeting stage then transfers the retrieved human motion to the new scene while preserving bone lengths, foot grounding, and temporal smoothness. The retargeted joint trajectories serve as sparse control signals for a controllable masked motion generation model, which synthesizes dense, smooth full-body motion consistent with the established interaction structure. InterAssembly is, to our knowledge, the first system to generate coherent, long-horizon, multi-object HOI activities, including tool use, object transfer, and multi-step tabletop cooking, at durations far beyond existing HOI generation methods, with low human-object penetration and consistent scene evolution throughout.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.