acceptodds
Under review as a conference paper at ICLR 2027

Can VLMs Judge Motion? From Capability Diagnosis to Reliable Deployment

Abstract

Mainstream evaluation metrics for text-driven human motion generation (T2M) compute global similarity scores, offering no fine-grained assessment of textmotion alignment. Vision-language models (VLMs), with their capacity for multimodal semantic understanding, are a natural candidate to fill this gap. However, 3D human motion involves spatial depth, skeletal topology, and temporal continuity that lie beyond typical 2D visual reasoning, leaving VLM reliability in this domain unverified—and deploying an unverified evaluator risks producing misleading conclusions in practice. To investigate this, we present a four-part study. (1) We construct FineMo (Fine-Grained Motion Semantic Alignment Benchmark), a diagnostic benchmark comprising 14,400 questions across six data-driven semantic dimensions, three alignment levels, and three progressive task types. (2) We evaluate six mainstream VLMs on FineMo, revealing three failure modes that persist consistently across all models and tasks: Depth Perception Failure, Topology Discrimination Failure, and Temporal Resolution Failure. (3) Tracing each failure, we identify two root causes—rendering-stage information loss and model-stage processing limitation—each verified through controlled experiments. (4) Building on this analysis, we propose an adaptive evaluation framework that diagnoses failure modes from text semantics and activates targeted interventions accordingly, achieving 57% overall accuracy—16 points above baseline and 9 above blanket intervention—without modifying the model. Together, these contributions provide a systematic characterization of VLM capability boundaries in motion evaluation and an evidence-based framework for their reliable deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.