When Should a World Model Switch? Surprise-Gated State-Space Models for Hierarchical Events
Abstract
A predictive world model faces a recurrent dilemma: preserve context while an event is stable, yet rapidly reorganize state when the current dynamics no longer apply. Most video models leave this decision implicit, and most event-boundary detectors mark changes without letting those changes alter the model's internal dynamics. We introduce , a multimodal latent world model that operationalizes Event Segmentation Theory as explicit predictive computation: becomes an endogenous control signal, and a learned hierarchical jointly (i) expands a selective state-space model's discretization step, attenuating stale history, and (ii) retrieves an input-conditioned that initializes the next state. Counterfactual utility regularization teaches the gate to switch only when reorganization improves future prediction, and a learned prior supplies event hazards during open-loop rollout; we prove a monotone bound on the direct contribution of pre-boundary state that a fully open schema-mixture removes exactly. improves controlled open-loop prediction by ( seeds, ) (OOD – in seeds across all four OOD families); on Kinetics-GEBD its internal gate reaches Rel. Dis. mean-F1 (zero-shot 5-seed ensemble on val clips with extracted features; per-seed ) without boundary-supervised representation learning, and the same gate reduces open-loop prediction error on KTH Actions at recurring action-regime starts while remaining a detector at exogenous cuts. In naturalistic fMRI, endogenous innovation adds unique cross-validated predictivity of (; subjects) in a pre-specified event network beyond audiovisual features and raw innovation, replicating on an independent film (, ), while the gate's discrete boundaries dissociably track sensory-change cortex ( subjects). Together, these results establish as a principled mechanism for continuous world modeling, and show that its functional decomposition reflects the organization of the human brain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.