acceptodds
Under review as a conference paper at ICLR 2027

IntentAnimator: Human Animation from Compositional and Underspecified Motion Intents

Abstract

Controllable human animation is commonly formulated as reproducing the motion of a driving video. However, real creative workflows often combine sparse pose keyframes, selected-joint trajectories, and local action descriptions, while specifying only selected time points, subjects, or body parts. We introduce Intent-Driven Human Animation, a new setting for synthesizing complete and coherent human videos from such compositional and underspecified motion intents. This setting requires jointly grounding heterogeneous constraints while plausibly completing the human dynamics they leave unspecified. We propose IntentAnimator, a unified framework comprising Omni-Intent Grounding (OIG) and Structure-of-Thought (SoT). OIG converts state, path, and text intents into a shared token sequence and uses Intent-RoPE to assign modality-appropriate positional anchors, enabling all intents to be jointly interpreted. Within diffusion, SoT supervises a dedicated structure-video stream in a human body-part domain to progressively complete missing dynamics and uses its evolving and semantically rich representations to guide pixel synthesis, improving motion coherence and reducing human deformation. We further construct a user-intent benchmark from annotations by professional creators and non-expert users, showing that natural motion intent is highly combined and sparse. Extensive experiments demonstrate that IntentAnimator improves intent following, motion plausibility, and visual quality over existing related baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.