Not All Regions Need Equal Tokens: Motion-Adaptive Representation Allocation for Complex Human Image Animation
Abstract
Complex human motion creates spatially non-uniform representation demands. Fast-changing body regions require finer representations, while slowly varying regions can be represented with fewer tokens. However, existing video Diffusion Transformers typically use a uniform patch grid, assigning the same token density to regions with substantially different representation needs. We rethink complex human image animation from this representation-allocation perspective and propose Motion-Adaptive Token Allocation (MATA), which reallocates spatial representation capacity according to motion demand. MATA assigns denser representations to motion-intensive regions while compressing redundant tokens in low-motion regions, allowing additional capacity to be directed where complex motion requires it without uniformly increasing the token budget. On HyperMotionX, MATA improves generation under complex human motion, reducing FVD by 7.97% and increasing by 6.54 percentage points over UniAnimate-DiT 5B, with only a 1.29% increase in estimated main-DiT FLOPs. These results highlight the importance of spatially non-uniform representation allocation for complex human animation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.