acceptodds
Under review as a conference paper at ICLR 2027

TrajWAM: Trajectory-Conditioned World Action Modeling for Vision–Language Driving

Abstract

World Action Models (WAMs) connect future scene prediction with action planning for autonomous driving. Multi-view video prediction requires camera-wise image decoding, while trajectory-agnostic bird's-eye-view (BEV) prediction leaves the influence of future ego motion implicit. Thus, we present TrajWAM, a trajectory-conditioned WAM that jointly learns BEV future representations and a vision-language driving policy. TrajWAM links world prediction and planning through a shared backbone by conditioning DINOv3-aligned BEV predictions on ego trajectories. Training combines reconstruction of observed futures with a feature-separation objective that encourages predictions to respond to trajectory changes, without requiring future labels for perturbed trajectories. In inference, the planner directly generates waypoints while bypassing the auxiliary future-prediction branch. On all 220 Bench2Drive routes, TrajWAM achieves 91.53 Driving Score and 75.45% Success Rate, exceeding a schedule-matched planning-only fine-tuning baseline by 5.31 Driving Score points and eight successful routes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.