acceptodds
Under review as a conference paper at ICLR 2027

PhysWAM: Physically Consistent World–Action Modeling for Autonomous Driving

Abstract

World–action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world–action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP). CPP unprojects generated depth, transforms the resulting points using generated SE(3) ego motion, and penalizes their distance from LiDAR points transformed using recorded ego motion. This geometric loss promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, a simple label-free consensus rule selects among sampled trajectories without a learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers robustly to unseen driving environments without additional training. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.