XTrajectory: Learning World Dynamics in Vision-Language Models through Unified Future Trajectory Prediction
Abstract
Future prediction is essential for agents that reason and act in the physical world, yet existing visual world models commonly represent future dynamics through pixels, latent states, or specialized motion modules. We investigate whether predictive physical dynamics can instead be acquired directly by a general-purpose vision-language model (VLM) through structured future trajectory prediction. We introduce XTrajectory, a unified framework that represents diverse future motion as autoregressively generated coordinate sequences through the VLM's native language interface, without introducing an additional trajectory decoder or dynamics head. XTrajectory unifies the future motion of visual points, physical objects, human hands, and robot end-effectors under a common history-to-future formulation, and learns this capability through geometrically verifiable reinforcement post-training. Experiments on existing trajectory benchmarks and our XTrajectory-Bench show improvements in future trajectory forecasting and strong transfer to unseen manipulation and physical-interaction scenarios. Evaluation on general multimodal benchmarks further reveals meaningful, though non-uniform, transfer beyond trajectory prediction. These results suggest that future trajectories provide a compact, interpretable, and directly verifiable interface for introducing predictive world dynamics into general-purpose VLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.