HumanMoveVQA: Can VideoMLLMs reason about Human Movement from Videos
Abstract
Multimodal Large Language Models (MLLMs) can describe what a person does in a video, but not where they go. Answering “does the person end up left of where they started?” or “how many times do they turn clockwise?” requires reasoning about global trajectory and orientation in 3D over time, beyond apparent motion in the image, which existing benchmarks, focused on scene-level events or local joint articulation, do not probe. We introduce HumanMoveVQA, a benchmark for reasoning about global human trajectory and orientation in exocentric video. We propose a scalable multi-stage pipeline that lifts 2D videos from five human-motion datasets into world-consistent 3D SMPL-X motion tracks anchored to the person’s first-frame position and heading, and converts them into discrete motion events. From these, we generate 12,282 multiple-choice questions over 127 test videos (15–90 s) across seven categories spanning event detection, aggregation, ordering, and trajectory-level inference. The best of 12 Video MLLMs reaches a chance-normalised score of 14.0, against 78.7 for humans. Fine-tuning an open-source model on our world-consistent supervision roughly triples its score, showing that the problem is learnable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.