Right Ranking, Wrong Reasons: Closed-Loop Causal-Channel Falsification for World-Model Evaluation
Abstract
World models are increasingly evaluated by whether their closed-loop rollouts recover the values or ranking of downstream policies. We show that even exact value and ranking agreement can be causally misleading. An error in the observation channel changes the actions induced by an observation-based policy, while an error in the physical transition model changes the consequences of those actions; the two errors can cancel. We give a finite constructive example in which two policies receive the exact correct values and strict ranking even though both policies take the opposite initial action and the learned physical model swaps the consequences of the two actions. We then introduce Closed-Loop Causal-Channel Falsification for World Models (C3F-WM), a four-corner audit that separates observation-mediated action error, physical-model error, and their cancellation. The empirical protocol combines detector-independent component endpoints, practical error floors, tie-safe calibration ranks, and a continuous train-only support distance. In an independent MetaDrive replication with disjoint train, calibration, and test seeds, the aligned action/support detector is evaluated on 64 units (31 positive, 33 negative). Its area under the receiver operating characteristic curve (AUROC) is 0.999, compared with 0.849 for action-only and 0.911 for physical-only. Seed-cluster bootstrap 95% lower bounds for the two AUROC margins are 0.079 and 0.017. At the calibration-selected threshold, recall is 0.903 at zero false-positive rate. These results demonstrate complementary detection of observation-mediated and physical errors. On 64 held-out scenario clusters, a separately trained three-initialization red-green-blue (RGB) ensemble obtains 2.832/2.828 m current/next gap mean absolute error (MAE) and 0.981/0.981 lead-presence accuracy, demonstrating learned recovery of audit-relevant driving semantics. Together, these findings connect policy-ranking evaluation with channel-specific diagnosis of world-model errors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.