Verifiable 4D World-Action Model for Robust Robot Control
Abstract
World-action models couple imagined visual futures with robot actions, but RGB-only imagination requires recovering geometry and motion from appearance that can change across scenes. We propose S4NDBOX, a 4D world-action model that jointly predicts RGB, depth, view-specific geometric motion fields, and robot actions, explicitly connecting visual evolution to 3D robot motion. The explicit 4D future supports action generation and candidate search based on geometric agreement between predicted futures and paired actions before execution. Across both simulated and real-world distribution shifts, S4NDBOX improves robustness over RGB-only control. On RoboTwin, after training only on clean scenes, S4NDBOX improves success in the hard (domain-randomized) setting by 13.3 percentage points over the RGB-only baseline. We further evaluate zero-shot simulation-to-real transfer on Stack Cups, training on simulated demonstrations and deploying on a physical robot without real-world task demonstrations or fine-tuning. Under additional shifts in object pose, lighting, and clutter, with camera-pose shifts additionally evaluated in the single-view setting, our 4D model improves the mean task score over RGB-only by 41.7 and 40.0 points in three-view and single-view settings, respectively. Website: https://demo1-indol-gamma.vercel.app/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.