CCPL: Continuous Causal Predictive Learning for Spatially Queryable Visual Representations
Abstract
We introduce Continuous Causal Predictive Learning (CCPL), a self-supervised framework that accumulates a causal state from local glimpses and predicts a continuous-valued next-glimpse latent conditioned on real-valued relative pan and zoom through a conditional flow matching objective. Unlike masked or autoregressive image models, whose predictions are specified by a mask position on a fixed patch grid or a fixed scan order, CCPL conditions each prediction on a continuous relative move, and hence the frozen model can be queried at independently specified positions and scales. After pretraining on ImageNet, query error decreases as a prescribed random-walk history grows and increases when recurrent memory is reset or query-move conditioning is shuffled. These controls provide operational evidence that the causal state accumulates information useful for spatial prediction across the sequence. We then freeze the observer and predictor and train an explorer with a category-label-free terminal reward given by normalized query error from the final feature. At a fixed budget of 1024 glimpses, learned scans reach the lowest terminal query error among all evaluated scans on both CCPL models and raise final-feature linear-probe top-1 over matched random walks by 1.18 and 1.34 points. However, recognition-leading lattice scans have higher query MSE than random walks. On the learned-scan's final and mean features, two-layer probes raise top-1 by 9 to 11 points over the corresponding linear probes, suggesting that class information in CCPL features is better accessed nonlinearly. Thus, a frozen predictive model can supply a category-label-free acquisition signal, but reconstruction and recognition favor different acquisition strategies and require separate evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.