LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue
Abstract
Large multimodal models (LMMs) describe dynamic scenes, but grounding successive questions in an entity’s changing 3D position requires a representation that connects dialogue references to physical time. We formulate trajectory-grounded multi-turn spatiotemporal dialogue, where a model answers questions and returns metric positions over the requested interval. We introduce Track4D-Bench to associate these dialogue references with calibrated observations and timestamped trajectories across 526 clips, 23.5k frames, and 7.5k entity tracks. To retain the queried entity throughout a conversation, we propose LMM-Track4D with a geometry-conditioned target state. This state preserves entity context across turns and selects spatial and temporal evidence for recovering the entity’s trajectory. The resulting representation separates dialogue order from the physical times being queried, allowing later questions to revisit earlier intervals. On Track4D-Bench, LMM-Track4D improves CIDEr by 0.247 over the strongest same-data supervised fine-tuning baseline, alongside gains of 10.2 and 10.5 points in BLEU-4 and METEOR. Geometric evaluation shows a 19.4 percentage-point gain in trajectory-step accuracy over direct coordinate regression. External evaluations further assess dynamic-scene question answering and language-conditioned grounding. These results support geometry-conditioned target states for connecting persistent entity references with metric trajectory reasoning in LMMs. Our code and dataset will be publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.