Auditing Misalignment in Multi-Modal World Modeling
Abstract
An omnimodal world model can predict one physical outcome while generating another. We distinguish internal alignment with the model’s own forecast from external alignment with a physical reference, using matched predictive contracts over events, magnitudes, and, when available, timing. Two intervention ladders progressively condition generation on either the model’s own contract (A) or the physical contract (B), ranging from scene-only prompts to event, magnitude, and keyframe instructions. In a matched Super push cohort, 39/40 jointly valid pairs have correct textual forecasts but incorrect measured video endpoints; other source-bound cases exhibit agreement on an incorrect physical observable. Increasingly explicit conditioning changes whether the prescribed motion occurs, but does not establish reliable quantitative repair. We further test whether these discrepancies matter for downstream learning under fixed-budget supervision. Screening actions using the model’s own textual forecasts increases push-planning success from 0.338 to 0.635. At matched retained action support, gated supervision also reduces mean label and fitted-grid errors, although additional task-success gains over the 0.615 control remain unresolved. Together, these results provide a framework for auditing physical world models across forecasting, generation, and downstream learning, while separating three claims that are often conflated: cross-modal agreement, successful physical repair, and downstream utility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.