When Do World Models Recover Latent Factors?
Abstract
World models can mix latent variables while task predictors compensate, preserving predictive accuracy. The existing low-degree guarantee requires direct coordinate readouts to remain first-order. We ask when minimizing the average Fourier degree of predictors sharing a representation recovers coordinates without this separate constraint. For cumulative task families defined by maximum interaction order, we completely characterize coordinate identifiability among reversible encodings of d binary variables under generic continuous and uniform finite-output task laws. Population minimizers recover coordinates, up to permutations and independent sign flips, if and only if the selected task spaces’ joint projection kernels distinguish pairs differing in one variable from all other distinct pairs. Cumulative cubic tasks suffice for d ≥ 5; every failing selection admits an explicit reversible mixed encoding. When each variable takes q ≥ 3 values, every proper, nonconstant cumulative family identifies coordinates. We also prove approximate recovery guarantees and give a reconstruction algorithm running in time polynomial in the state count n = 2^d. Readouts supervised on all states show that fixed representations from frozen Transformers support coordinate recovery through task kernels. After graph correction, cumulative cubic readouts achieve 100% coordinate accuracy for d = 6–11, while the pipeline rejects all tested non-identifying controls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.