acceptodds
Under review as a conference paper at ICLR 2027

STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation

Abstract

Spatial-temporal prediction and geometric grounding are increasingly critical for high-precision robotic manipulation, yet remain challenging for world action models (WAMs). While these models capture temporal context through future prediction, RGB-only modeling does not explicitly represent the 3D geometry required for precise control. Consequently, visually plausible futures may still fail to expose action-relevant spatial constraints, limiting their ability to improve action generation. We propose STARRY, a world action model that aligns future spatial-temporal prediction with action generation through a unified diffusion process, enabling actions to be informed by the predicted evolution of scene appearance, spatial geometry, and end-effector motion. Meanwhile, to make more effective use of the predicted spatial-temporal information for action generation, we introduce Geometry-Aware Selective Attention Modulation (GASAM), which converts geometric cues derived from the predicted spatial-temporal variables into token-aligned weights, selectively directing action attention toward regions relevant to future interactions. On RoboTwin 2.0, STARRY achieves 93.82% / 93.30% average success under Clean and Randomized settings across 50 bimanual tasks. Real-world experiments show that STARRY improves average success from 35.0% to 67.1% across four evaluation batches compared with . These results demonstrate that spatial-temporal action-centric world modeling improves robotic manipulation by aligning future spatial-temporal prediction with action generation and selectively guiding action attention with predicted geometry.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.