acceptodds
Under review as a conference paper at ICLR 2027

CAT2M: Camera-Aware Text to Motion Generation

Abstract

Recent world models have achieved notable progress in synthesizing realistic videos with controllable camera movements; however, achieving fine-grained human motion control under dynamic viewpoints remains a significant challenge. Meanwhile, existing text-to-motion methods typically assume fixed camera settings, which limits their applicability in world model settings where human motion should be consistent with the text description under changing camera viewpoints. To address this gap, we introduce a new task: camera-aware text-to-motion generation, where human motion is generated based on both text descriptions and given camera trajectories. To facilitate this task, we construct the CAT-Motion dataset, a large-scale dataset with over 100K samples of paired human motion and camera trajectories, each annotated with rich captions including viewpoint-specific motion descriptions. Based on this dataset, we propose CAT2M, a unified framework for motion generation with text and camera trajectory conditions. A shared coordinate frame anchored to the first camera preserves human-camera relative geometry, while temporal camera cross-attention conditions generation on the changing viewpoint. We further introduce differentiable screen-space supervision through a cumulative training scheme that retains diffusion and 3D joint objectives, constraining the projected appearance alongside the underlying 3D motion. Extensive experiments demonstrate that our approach outperforms existing baselines, showing strong ability in producing human motions that are semantically accurate and geometrically consistent under camera constraints and text descriptions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.