acceptodds
Under review as a conference paper at ICLR 2027

World Turing Machine: Learning Read–Write Memory for 3D Worlds

Abstract

Spatial world modeling requires collecting evidence across viewpoints and maintaining it in a persistent state for predictions at unobserved views. However, organizing and accessing such accumulated information over time remains a central challenge. In this study, we introduce the World Turing Machine (WTM), which addresses this through learned read–write computation over an Address–Content memory. This memory evolves with observations while the learned read–write operations remain fixed at inference. Before each write, the Reader predicts incoming visual features from the current memory and the condition. The Writer reuses the Reader's access weights to correct Content with prediction residuals and update Address with observation-derived geometric evidence. Local and global predictive objectives train acquisition, retention, and inference across viewpoints without directly supervising memory. The same Reader supports forward prediction at unvisited poses and inverse pose matching against target-frame features, without an additional localization head. On ScanNet, WTM outperforms evaluated novel-view synthesis and feed-forward 3D reconstruction baselines in held-out feature prediction. With a simple search over candidate poses, WTM achieves 65.7% success on ViewSuite tasks without task-specific policy training, surpassing ViewAgent by +13.1 percentage points. Analyses reveal emergent spatial organization during online memory formation and show that the learned read–write mechanism generalizes across 3D environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.