GEAR: Geometry-Aware Rewards for Human Motion in Text-to-Video Generation
Abstract
Reinforcement learning has shown strong potential for improving text-to-video (T2V) generation, yet generated human motion remains prone to structural deformations and temporal inconsistencies. % Existing post-training rewards provide incomplete supervision for this problem. Holistic learned scores entangle body correctness with other visual factors, while human-centric methods rely on either coarse structural cues or a single evidence representation. Moreover, optimizing structural signals alone can favor low-motion outputs, compromising prompted action execution. % To address these limitations, we propose GEAR (Geometry-Aware Rewards), an RL post-training framework that jointly addresses structural credit assignment and action preservation. GEAR combines direct 2D keypoint measurements, which retain image-space evidence of structural abnormalities, with reconstructed 3D motion evidence that captures depth- and orientation-aware information. It further introduces a requirement-level semantic regularizer that evaluates decomposed motion requirements without fixed phase-to-time correspondence, discouraging low-motion and incomplete-action solutions. % Experiments across different T2V backbones show that GEAR reduces human deformation while preserving prompted action execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.