MineLoop: Inner-Time Computation for Autoregressive Visual World Models
Abstract
Autoregressive visual world models turn each sampled token into evidence for later predictions. Standard feed-forward predictors sample each token after a single pass, without an explicit stage for refining uncertain predictions under the same history. We introduce MineLoop for inner-time world modeling, giving each prediction an internal trajectory before it enters the world history. MineLoop encodes the committed visual-action prefix as causal evidence, initializes a target-indexed future state, and recurrently updates it with a shared prediction core while the evidence remains fixed. Only the terminal state emits a visual token. On action-conditioned Minecraft rollout, MineLoop improves FVD, LPIPS, SSIM, and PSNR against MineWorld controls at Small, Medium, and Large model tiers, with FVD reductions of 16.3%, 10.5%, and 7.3%, respectively. Counterfactual camera-action tests show stronger action alignment, and an inverse-dynamics evaluator recovers conditioning actions with comparable macro-F1. Across 64 held-out sources from the Small model, 23.13% of visual tokens change from wrong at K1 to correct at K4, while 0.093% change in the opposite direction. The advantage over a standard autoregressive model is concentrated in motion and occlusion regions. Continuing beyond the trained terminal step raises confidence while reducing accuracy, showing that MineLoop learns a finite-depth prediction program with a specific terminal state. These results identify token commitment as a productive locus for recurrent computation in visual world models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.