Can VLMs Judge Motion? From Capability Diagnosis to Reliable Deployment
Abstract
Mainstream evaluation metrics for text-driven human motion generation (T2M) compute global similarity scores, offering no fine-grained assessment of textmotion alignment. Vision-language models (VLMs), with their capacity for multimodal semantic understanding, are a natural candidate to fill this gap. However, 3D human motion involves spatial depth, skeletal topology, and temporal continuity that lie beyond typical 2D visual reasoning, leaving VLM reliability in this domain unverified—and deploying an unverified evaluator risks producing misleading conclusions in practice. To investigate this, we present a four-part study. (1) We construct FineMo (Fine-Grained Motion Semantic Alignment Benchmark), a diagnostic benchmark comprising 14,400 questions across six data-driven semantic dimensions, three alignment levels, and three progressive task types. (2) We evaluate six mainstream VLMs on FineMo, revealing three failure modes that persist consistently across all models and tasks: Depth Perception Failure, Topology Discrimination Failure, and Temporal Resolution Failure. (3) Tracing each failure, we identify two root causes—rendering-stage information loss and model-stage processing limitation—each verified through controlled experiments. (4) Building on this analysis, we propose an adaptive evaluation framework that diagnoses failure modes from text semantics and activates targeted interventions accordingly, achieving 57% overall accuracy—16 points above baseline and 9 above blanket intervention—without modifying the model. Together, these contributions provide a systematic characterization of VLM capability boundaries in motion evaluation and an evidence-based framework for their reliable deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.