acceptodds
Under review as a conference paper at ICLR 2027

GeoScale: Inference-Time Scaling of Geometry-Aware Video World Models for Robot Manipulation

Abstract

Video world models provide a promising interface for robot manipulation by predicting future visual states that can be converted into actions. Geometric consistency and motion plausibility in these predictions are critical for recovering executable action trajectories, as robots must carry out these actions in the physical world. We observe that different random seeds at inference time produce rollouts with substantially different geometric quality, suggesting that better predictions can be obtained by searching over seeds. We therefore propose **GeoScale**, an inference-time scaling framework that searches for seeds yielding geometrically stronger rollouts. GeoScale uses feed-forward 4D reconstruction as geometric feedback for both candidate selection and guidance during denoising. It progressively scores and prunes candidate rollouts to concentrate computation on promising candidates, while applying differentiable geometric guidance to refine the surviving candidates at selected denoising steps. Experiments on real-world Droid and simulated RLBench show consistent gains in depth accuracy, 3D reconstruction, and point correspondence while preserving visual quality. GeoScale also improves human-rated task completion and simulator execution success, demonstrating more actionable rollouts for robot manipulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.