Programmatic Search Policies For Planning
Abstract
We study the problem of transferring knowledge from solved embodied planning problems to related future problems. Embodied planning problems are typically solved via bi-level planning, where a high-level planner explores an environment by proposing sub-goals, which are then realized by a low-level control system. We are interested in the setting where the space of sub-goals must be rich enough to encode useful sub-goals and where the low-level control system cannot realize all sub-goals; this setting makes the low-level control system the dominant cost during exploration. Thus, the key challenge is to propose useful sub-goals that allow the planner to efficiently find trajectories that maximize reward. We refer to a system that generates sub-goals during exploration as a search policy. We describe an algorithm that synthesizes programmatic search policies — programs that generate candidate sub-goals for a given state, which drive the planning process. The search policies are represented as Python programs, which allows them to generate sub-goals in an expressive language and delegate some computation to the program interpreter, and they are amenable to being efficiently updated to novel environments via Language Models. Thus, by synthesizing a programmatic search policy, our proposed algorithm stores the knowledge acquired during the exploration in a representation that allows easy adaptation to novel tasks. We empirically show, in simulated environments, that the policies synthesized by our algorithm can be used to transfer knowledge across tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.