Task-specific Context Optimization for Training-free Robot Planning
Abstract
Frontier vision-language models (VLMs) can serve as off-the-shelf high-level planners in hierarchical robot policies, using multimodal reasoning to infer subtasks for low-level execution. However, their out-of-the-box use leaves ambiguity in offline subtask annotation (for low-level policy fine-tuning) and online subtask planning, resulting in inconsistent guidance that we find degrade long-horizon task performance. We introduce TAsk-Specific Context Optimization (TASCO), a framework that adapts frozen frontier VLMs through task-specific context to produce consistent offline annotations and align causal online predictions with them. Specifically, TASCO queries a separate VLM, termed the context optimizer, to induce this context from a reference demonstration, providing a common basis for consistent task decomposition. Next, for temporally consistent online planning, the context optimizer iteratively refines this context by diagnosing mismatches between offline annotations and online predictions, accepting updates only when their consistency improves. Across multiple frontier VLMs and low-level policies, TASCO substantially improves complex, long-horizon task performance on simulated and real-robot benchmarks, while naive deployment yields surprisingly poor results. These findings establish TASCO as a model-agnostic framework for harnessing the growing reasoning capabilities of frontier VLMs by aligning high-level planning with low-level execution. Project page is available at https://iclr-tasco.github.io/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.