Locating Memory In Predictive World Models: Interventional Diagnosis Beyond Probing
Abstract
Linear probing is the default instrument for asking what a learned representation encodes, but it cannot say how a model retains that information. We exhibit three world models — a stateless Joint Embedding Predictive Architecture (JEPA), a GRU and an LSTM — that reach identical probe accuracy (1.000) on the same delayed working-memory task while implementing mechanistically incompatible forms of memory. To separate them we introduce nested interventional diagnostics: stable retention (SR) asks whether information about a past causal event survives a latent rollout; an interventional causal separation test, with a 2x2 extension that toggles recurrent state, asks where the memory resides; and the causal-subspace Lyapunov factor lambda_v(t) asks how the predictor sustains it. Three findings follow. Memory locus is a property of the trained model rather than of the architecture class: GRU and LSTM seeds trained identically span from near-pure state-based to predominantly dynamical solutions, a spread invisible to probe accuracy. A stateless predictor sustains memory by selective subspace preservation: at a prediction horizon of K = 5 it contracts sampled latent distances on average (kappa 0.36) while amplifying the causal direction 6.48x relative to a typical direction, a condition strictly weaker than the global norm preservation that unitary and orthogonal recurrent parameterisations enforce. And a controlled ablation that redistributes a trained encoder's causal weight across an orthonormal basis, holding predictor, task and training protocol fixed, drives stable retention from 0.458 to 0.107 and causal-subspace amplification from 14.2 to 4.3: encoder geometry causally gates whether dynamical memory is available at all.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.