RoboSTG: Training-Free Spatial-Temporal Guidance for Generative Robot Policies
Abstract
Pretrained generative robot policies provide expressive behavioral priors, yet their deployment-time behavior remains limited by the objectives encoded in training. The central challenge is not simply how to apply guidance at inference time, but how to convert manipulation-specific requirements into differentiable objectives that are compatible with generative action synthesis. We introduce RoboSTG, a training-free framework for reshaping frozen generative robot policies with manipulation-oriented spatial-temporal guidance. RoboSTG expresses task requirements through differentiable spatial fields and relational constraints in physical space, and lifts them to action-chunk objectives via substage-aware spatial-temporal modulation. These objectives produce gradients that directly steer iterative action generation, without policy fine-tuning, learned value models, or candidate ranking. Experiments on LIBERO and RoboTwin and four real-world manipulation settings show strong task success and robust task-aligned execution. Together, the results demonstrate that structured, temporally aligned manipulation objectives provide an effective interface for adapting pretrained generative policies at inference time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.