acceptodds
Under review as a conference paper at ICLR 2027

Sparse Trajectory Tokens for Metric Motion Queries in Multimodal Language Models

Abstract

Video-language models can recognize motion yet struggle to answer questions about metric displacement and motion rate when scale and camera motion must be inferred from images. We study an interface that supplies an MLLM with externally reconstructed 3D object and camera trajectories alongside video. Our representation associates sparse, part-aware keypoints with metric world-frame tracks, splits tracks at occlusions and motion events, and encodes adaptively fitted spline samples as trajectory tokens. A lightweight projector injects these tokens into Qwen3-VL-8B without modifying its vision encoder. We also construct TrajQA-B from Stereo4D tracks, with token-side answer checks and controls for text and option shortcuts. On its 7,500 four-choice numeric questions, the trajectory model obtains 82.2% accuracy, versus 48.3% for a separately fine-tuned frames-only reference without timestamps or guaranteed query frames. Replacing trajectories with those of another scene lowers accuracy to 29.6%; the model's generated metric evidence also tracks the reference positions and motion quantities. These results show dependence on supplied trajectories for these QA tasks. Caption and QA training exclude the evaluation scenes; source-video separation for caption training has not been audited. The benchmark uses reconstructed rather than independently measured motion references.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.