acceptodds
Under review as a conference paper at ICLR 2027

Solis: Streaming Unified World Models for Robotic Interaction

Abstract

Robotic interaction couples actions, changes in the environment, and progress toward a task. Modeling these modalities jointly offers a foundation for both generating actions toward desired outcomes and anticipating the consequences of actions. We present Solis, a streaming unified world model that learns a joint distribution over observations, actions, and progress-derived values using a single video diffusion transformer. Using different inputs as conditions, the same model can be used for simulation, planning, or evaluation. For practical deployment, closed-loop feedback is important for adapting to changes in the environment, yet restarting diffusion sampling after every observation makes frequent feedback costly. To address this, we employ temporal block-causal attention and natively pretrain our model with independently sampled chunk noise levels. At inference, we adopt a rolling diffusion schedule to maintain future video, action, and value chunks at progressively increasing noise levels: the leading chunk becomes executable while later chunks remain partially denoised. After execution, the horizon shifts with a new observation and a fresh noise chunk enters. This design raises feedback frequency without extra computational cost. With cross-embodiment pretraining and per-benchmark post-training, Solis achieves 98.4% average closed-loop success across four LIBERO suites and 88.8% and 89.1% across 12 RoboTwin 2.0 tasks in clean and randomized scenes. Furthermore, Solis shows strong future prediction and accurate value estimation for task evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.