World-as-Prompt: Rewriting Robot Observations for Robust Policy Execution
Abstract
Vision-language-action and world-action policies can acquire useful manipulation from a modest dataset in one visual domain, yet illumination, sensor noise, object color, and workspace texture turn the same task into an out-of-distribution (OOD) observation. Collecting action-labeled data for each such factor is slow and hard to scale, and wrapping the policy with an extra agent or corrector still leaves it reading the shifted image. We introduce **World-as-Prompt (WAP)**, in which an in-domain reference image and a structured appearance description form a prompt for the world the frozen policy should see. We instantiate that prompt as a world-to-world (W2W) translator. At each control step, the translator conditions on the depth of the current OOD observation, that reference, and the description, changing nuisance appearance while retaining geometry and interaction state. The module is attached before an existing vision-language-action model (VLA) or world-action model (WAM), and requires neither new action-labeled robot data nor policy fine-tuning at deployment. We study WAP on RoboTwin2.0 Clean2Random, LIBERO-Plus, and real-robot cloth folding. Across simulation and real-robot evaluations, WAP improves every evaluated VLA and WAM by to percentage points in OOD success, demonstrating that it restores much of the competence that frozen policies lose under visual shift and offers a plug-and-play way to deploy existing VLAs and WAMs in new visual conditions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.