Action-Guided In-Context Example Retrieval for Embodied Task Planning
Abstract
Given only a small set of annotated instruction-plan demonstrations, large language models (LLMs) have shown promising performance in embodied task planning through in-context learning. However, existing approaches primarily retrieve in-context examples based on linguistic or visual similarity, so the retrieved examples may fail to include the actions required by the query task and thus provide limited guidance for generating an appropriate task plan. To address this limitation, we propose an action-guided retrieval approach that retrieves examples relevant to the actions implied by the query to better guide task plan generation. To estimate this relevance, we perform action-guided representation learning based on action sets and action sequences derived from annotated demonstrations, yielding instruction embeddings that capture the action types and their action patterns necessary to complete the task. We further introduce an action-preserving augmentation strategy that diversifies training instructions while preserving their underlying action information to alleviate the scarcity of annotated demonstrations. Experiments on embodied task planning benchmarks show that the proposed method outperforms existing retrieval strategies when integrated with multiple LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.