Say Where: Language and Goal Sets for Frozen Visual Planning
Abstract
A task often accepts many outcomes, while a goal image depicts only one. We study an interface that preserves these alternatives for planning with a frozen visual world model. Task specifications define sets of acceptable states; finite clouds of their visual representations supply the planner’s targets. This supports joint constraints, alternative destinations, precision, and factors left free by the task. Across four simulated environments, a shared-start comparison with the same nearest-goal cost increases average success from 31.1% with one accepted target to 52.3% with a large cloud. Targeted studies show that jointly valid targets improve simultaneous constraint satisfaction, and retaining alternatives improves attainment over committing to the initially nearest branch. We then learn to construct goal clouds from language through retrieval or generation. On the full endpoint suite, the retrieval interface achieves 46.3% success versus 31.6% for one accepted image, and is especially effective on region and property instructions. Set-agreement measurements distinguish successful specification from remaining grounding errors. These results connect richer task specifications to useful planning behavior without retraining the predictive dynamics, while identifying precise constraints and selection among valid alternatives as continuing challenges.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.