acceptodds
Under review as a conference paper at ICLR 2027

Weak Models or Insensitive Metrics? Auditing and Training Fine-Grained Motion Captioning

Abstract

Motion captions are evaluated by n-gram overlap, embedding similarity, or an LLM judge that scores the caption as a whole, but whether these metrics notice one wrong motion fact, such as the acting side or the direction of movement, has not been tested attribute by attribute. Without that test, a weak model cannot be distinguished from an insensitive metric. We build MotionADV, a Chinese–English challenge set of caption pairs that differ in exactly one fact, from 494 clips whose captions were corrected by professional motion artists. Under a single wrong fact, n-gram and embedding-based metrics move less than a fifth of the way to an unrelated caption, and n-gram metrics even prefer the wrong caption to a fact-preserving paraphrase. A whole-caption LLM judge cannot both detect and localize the error. We define MotionFact, a fixed rubric of motion attributes over a typed action schema, and build on it MotionJudge, a decomposed evaluation protocol that checks each attribute against a reference with its own questions. With Qwen3.5-27B, MotionJudge reaches a detection rate of at least 0.90 and a pass rate of at least 0.89 on every attribute, and an ablation that adds the rubric and then the schema-based protocol shows what each contributes. We further train Motion-VLM, a general video-language model that reads multi-view renders and motion tokens, and states MotionFact facts before its caption. Under both judges, Motion-VLM outperforms five published motion captioners on HumanML3D and demonstrates superior cross-dataset transferability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.