PSWS: Learning Predictive Semantic World State for Multimodal Understanding
Abstract
Most multimodal understanding models optimize representations for recognizing the semantic label of the current observation. Such supervision, however, does not explicitly require the representation to preserve evidence that is predictive of subsequent semantic changes. We formalize this limitation as Semantic Transition Aliasing (STA): histories with the same current label but different next-label transitions may remain insufficiently separated in the learned state space. We propose Predictive Semantic World State (PSWS) model, a task-grounded and action-free predictive semantic-state model that infers a current state from past-only textual, acoustic, and visual observations and predicts its one-step evolution without accessing future observations at inference. PSWS factorizes current-semantic and transition-oriented evidence through dedicated latent queries, predicts modality- and interaction-specific state increments using an evidence-conditioned router, and aligns the predicted state with an EMA-encoded future target available only during training. A shared decoder grounds current and predicted states in the same label space. PSWS improves over reported current-recognition references and yields gains in future-label prediction and changed-transition performance under a unified causal protocol with supervision- and modality-matched baselines. Additional analyses evaluate transition separation, routing behavior, statistical reliability, and computational cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.