AW3D: Action Prototype Learning with World Modelling for Text-Driven 3D Human-Object Interaction Generation
Abstract
Realistic three-dimensional human–object interaction (HOI) animation is important for character animation, embodied simulation and humanoid control. Although existing text-conditioned HOI generators can produce physically plausible interactions, they remain limited in generating motions that faithfully follow textual instructions, often omitting required actions or executing them with incorrect timing or ordering. Here we present AW3D, an action-prototype framework with interaction world modelling, whose central idea is to explicitly organize high-level interaction intent into an ordered sequence of variable-duration action stages before dense motion generation. AW3D incorporates three key designs: Joint-Embedding Action Prototype Learning, which constructs a data-driven vocabulary of semantically meaningful interaction stages; Distribution-Guided Action Sequence Decomposition, which decomposes high-level intent into an ordered sequence of action prototypes and estimates their durations; and Interaction World Modelling for Role-Guided Dense Motion Generation, which adapts the resulting sequence to coordinate body, hand and object motions according to their roles in the interaction. Extensive experiments demonstrate that AW3D achieves state-of-the-art semantic alignment and motion quality, while remaining competitive in physical plausibility and interaction consistency. Code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.