acceptodds
Under review as a conference paper at ICLR 2027

GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation

Abstract

Vision-language-action (VLA) policies have shown strong performance in language-conditioned robotic manipulation, but standard action imitation provides limited supervision for learning explicit 3D geometry and future scene evolution. We propose GaussianDream, a feed-forward 3D Gaussian world model that introduces structured spatial and temporal supervision into VLA policy learning. GaussianDream inserts a set of learnable queries into the policy encoder to form a latent GaussianDream prefix. During training, the prefix is decoded by a current reconstruction head and a future prediction head to model the observed 3D scene and its short-horizon evolution. The current branch is supervised with RGB rendering and pseudo depth, while the future branch additionally uses future RGB, pseudo depth, and pseudo 3D scene-flow supervision. At inference time, all auxiliary reconstruction and prediction heads are removed, and only the learned GaussianDream prefix is retained for action generation, requiring no test-time Gaussian reconstruction or future rollout. GaussianDream achieves 98.4% success on LIBERO, 87.0% on LIBERO-Plus, and 54.8% on RoboCasa Human-50. We further validate GaussianDream on real robots using both single-arm and dual-arm platforms, achieving 50.0% average success on the single-arm tasks and 40.0% success on the higher-precision dual-arm manipulation task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.