acceptodds
Under review as a conference paper at ICLR 2027

Whose Motion Is It? Learning Motion-Source Attribution in Vision-Language Models

Abstract

Video-language reasoning requires understanding how visual content evolves over time. Yet such changes may arise from camera motion, target-object dynamics, or both, making their physical source ambiguous. We identify a recurring failure mode in current video-language models (VLMs), termed motion-source misattribution, where observed changes are attributed to the wrong physical source. Our empirical analysis shows that motion judgments are not reliably bound to the queried target, source information is recoverable but not reliably source-selective, and camera and object evidence interact non-additively when both motions co-occur. These findings reveal a clear gap: encoding motion cues does not guarantee organizing them by physical source. To address this gap, we propose MoSDeR (Motion-Source Decomposition and Reasoning), which structures temporal evidence into query-conditioned camera and object factors. MoSDeR first binds motion evidence to the queried entity, then routes the two factors into the native reasoning pathway, and finally composes them non-exclusively under concurrent motion. To evaluate motion-source attribution across diverse conditions, we curate a unified target-level benchmark spanning egocentric and allocentric settings, in-door interactions, autonomous-driving scenes, and diverse camera–object motion configurations. Across three VLMs, MoSDeR exceeds the runner-up in Overall accuracy by 3.64% on our benchmark and 6.75% on OmniVCHall on average, indicating consistent attribution gains and broader downstream benefits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.