Do Driving World-Action Models Imagine What They Plan? Image–Action Consistency for Evaluation and Learning
Abstract
Driving world-action models (WAMs) pair a native action with an imagined future. Interpreting these outputs as one coherent prediction requires evidence that they describe compatible motion and respond together to changes in conditioning. We introduce Image-Action Consistency (IAC), which evaluates this relationship through candidate-blind visual measurement and explicit observation coverage. IAC maps video and action into a shared space of yaw and interval longitudinal progress. The Alignment Score (AS) measures within-pair motion agreement, the Response Consistency Score (RCS) checks paired yaw responses to controlled condition changes, and the Grounding Score (GS) measures geometric agreement with independently logged futures. On held-out logged videos, visual progress estimates achieve tolerance accuracy on supported intervals, with coverage. Under the tested DriveWAM intervention protocol, response directions agree in of pairs with detectable responses in both outputs, yet only of all valid pairs produce material action responses. This gap shows why conditional agreement must be interpreted together with response coverage. Paired visual-motion targets also improve PDMS over matched-budget action-only fine-tuning by points on DriveWAM and point on WorldDrive. IAC thus provides both a diagnostic of joint predictions and a source of motion supervision for action learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.