acceptodds
Under review as a conference paper at ICLR 2027

XTrajectory: Learning World Dynamics in Vision-Language Models through Unified Future Trajectory Prediction

Abstract

Future prediction is essential for agents that reason and act in the physical world, yet existing visual world models commonly represent future dynamics through pixels, latent states, or specialized motion modules. We investigate whether predictive physical dynamics can instead be acquired directly by a general-purpose vision-language model (VLM) through structured future trajectory prediction. We introduce XTrajectory, a unified framework that represents diverse future motion as autoregressively generated coordinate sequences through the VLM's native language interface, without introducing an additional trajectory decoder or dynamics head. XTrajectory unifies the future motion of visual points, physical objects, human hands, and robot end-effectors under a common history-to-future formulation, and learns this capability through geometrically verifiable reinforcement post-training. Experiments on existing trajectory benchmarks and our XTrajectory-Bench show improvements in future trajectory forecasting and strong transfer to unseen manipulation and physical-interaction scenarios. Evaluation on general multimodal benchmarks further reveals meaningful, though non-uniform, transfer beyond trajectory prediction. These results suggest that future trajectories provide a compact, interpretable, and directly verifiable interface for introducing predictive world dynamics into general-purpose VLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.