Cognitive Liquid World Models: Continuous-Time Causal Verification over Decoupled Latent Manifolds
Abstract
Monolithic video world models compress multimodal observations into uniform representations updated over rigid discrete intervals (Δt=1.0). For embodied visual reasoning, this introduces four structural bottlenecks: categorical semantics dominate subtle spatial mechanics, rigid integer clocks miss irregular multi-scale transitions, static visual shortcuts bypass physical causal directionality, and competitive softmax normalization forces overconfident false positives under out-of-distribution inputs. We propose the Cognitive Liquid World Model (CLWM), reformulating physical world modeling as continuous-time causal hypothesis verification across a two-stage developmental firewall (an autonomous physics engine coupled to a frozen epistemic probe). CLWM resolves these limitations through four coordinated mechanisms: (1) asymmetric manifold bifurcation, decoupling an expressive semantic stream from an algebraically constricted geometry bottleneck stabilized by non-gradient continuous plastic codebooks; (2) top-down elastic ordinary differential equations (ODEs), where effective integration timescales warp dynamically across interaction boundaries to anchor persistent state transitions; (3) temporal group permutation optimization (TGPO), regularizing chronologically scrambled sequences to maximum Shannon entropy to enforce physical non-commutativity (A -> B = B -> A) and induce an emergent [0.40, 0.60] epistemic abstention attractor; and (4) an unconstrained continuous-time epistemic margin (CTEM) protocol that replaces closed-simplex probability competition with calibrated hypothesis verification. On Something-Something-v2, CLWM attains 0.9331 ROC-AUC (90.58% specificity) across all 174 classes and 0.9592 ROC-AUC on asymmetric causal transitions; targeted knock-outs collapse discriminability to chance baselines (0.4984, 0.5941, and 0.5546). Under a 50% frame drop (Δt=2.0), CLWM retains 99.35% of baseline discriminability without retraining, while prefix evaluations demonstrate sharp verification commitment at the physical contact boundary. Across 174 concurrent candidates, the unconstrained CTEM protocol yields a +0.2898 belief separation without competitive softmax normalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.