acceptodds
Under review as a conference paper at ICLR 2027

T4D: Comprehending the Dynamic World with MLLMs

Abstract

Multimodal large language models (MLLMs) have made rapid progress on 3D spatial reasoning, yet these successes remain confined to a static world. Real scenes evolve, and understanding them requires reasoning jointly over space and time—4D spatial reasoning. Existing benchmarks judge spatial change only qualitatively and leave dynamic egocentric–allocentric transformation untested; at comparable scale, open-source and expert models remain weak on dynamic scenes. We present T4D, a systematic study of dynamic spatial intelligence. We revisit spatial intelligence and, for the first time, construct the taxonomy of this field from a dynamic perspective. This yields six dimensions, instantiated as T4D-Bench: 10 tasks and 3,250 expert-verified QA pairs. Two difficulties stand in the way of scale: the clips do not share one 4D annotation—outdoor video supplies boxes, tracks, and ego poses, while egocentric video needs another geometric readout—and a language reference is ambiguous among similar objects. Our automated pipeline computes each answer from the native measurements and marks the queried entity on the frame, producing the benchmark and about 600K training QA pairs. On the algorithm side, (1) we propose the T4D network. It fuses explicit 3D point trajectories of instances with vision, and it attains state-of-the-art on T4D-Bench among MLLMs of comparable scale. (2) Trajectory-aligned training (TAP), a training recipe, supervises those instance trajectories in a fixed reference frame, aligning 2D vision tokens with the trajectory tokens and improving dynamic metric reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.