PHYFLOW: PHYSICS-AWARE RECTIFIED FLOW WITH FINE-GRAINED SEMANTIC ALIGNMENT FOR TEXT- TO-MOTION GENERATION
Abstract
Text-to-motion generation strives to synthesize human motion sequences that are semantically matched natural-language descriptions while remaining realistic, diverse, and physically plausible. Recent latent diffusion approaches, such as ARDT, attained competitive results by generating motion within a compressed latent space and refining realism through adversarial training and temporal semantic contrastive learning. However, there are still some limitations in their semantic conditioning as it relies on a single pooled CLIP sentence embedding, which discards the phrase-level structure of the prompt before it reaches the generator. Additionally, the alignment objective operates at sequence granularity, matching one sentence to one motion, and therefore supervises which actions occur but not when they occur. Moreover, diffusion-based sampling requires a large number of denoising steps at inference. In addition, generated motions frequently exhibit kinematic artifacts, including foot sliding, bone-length drift, infeasible joint configurations, and temporal jitter. We propose PhysFlow, a semantic- and physicsaware rectified-flow framework for text-to-motion generation. PhysFlow retains the two-stage latent generation paradigm while redesigning the text-conditioned generator. In the first stage, a VAE motion sequences into a compact continuous latent representation and reconstructs them to motion space. In the next stage, latent diffusion is replaced by a Rectified Flow Transformer that learns a direct velocity field from the noise distribution to the motion latent distribution, permitting inference with substantially fewer sampling steps. Semantic conditioning is provided by a frozen Qwen3 text encoder, which yields token-level contextual representations in place of a single global embedding. An Action Parsing Module further extracts action phrases, body-part cues, directional modifiers, and temporal relations from the input text; the resulting structured semantic tokens are injected into the generator through cross-attention. To strengthen temporal grounding, we introduce Fine-Grained Semantic Alignment, which aligns action-level semantic tokens with latent motion frames rather than aligning a complete sentence with a complete sequence. Beyond semantic modeling, PhysFlow incorporates physical and distributional regularization. A Physics Guidance Module imposes differentiable penalties on foot-contact inconsistency, bone-length deviation, joint-limit violation, and temporal non-smoothness, thereby constraining generated motions toward kinematic stability. To narrow the distribution gap without reinstating adversarial training in the generator stage, we adopt non-adversarial Distribution Matching objectives comprising moment and maximum mean discrepancy regularization at both the latent level and the evaluator-feature level. The complete training objective combines rectified-flow matching, fine-grained semantic alignment, physics guidance, and distribution matching, and thus avoids the optimization instability associated with adversarial latent diffusion training. We evaluate PhysFlow on the HumanML3D and KIT-ML benchmarks. On HumanML3D, PhysFlow improves semantic retrieval relative to ARDT, attaining Top-1/Top-2/Top-3 R-precision of 0.567/0.756/0.842 and an MM-Dist of 2.802. On KIT-ML, it improves motion distribution quality, reducing FID from 0.334 to 0.278 while maintaining a diversity score close to that of real motion data. These results indicate that token-level language representation and action-level temporal alignment account for the gains in semantic fidelity, whereas physics guidance and distribution matching account for the improvement in distributional quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.