Learning to Decompose and Act: Landmark-Based Manager–Worker Policies for Generalized Planning at Scale
Abstract
General policies learn from small instances to solve an entire instance class of a classical planning domain. Yet, as instances grow, so do the plans required to solve them, as well as the cost of evaluating alternatives at each decision. We introduce Landmarker, a manager–worker hierarchy that decomposes tasks into atomic subproblems: the manager generates one atom from a landmark vocabulary as the next subgoal, and the worker generates the primitive actions that solve it. The two policies are trained independently of each other with the help of polynomial width-based search: the manager via reinforcement learning, for which the search achieves its subgoals, and the worker from the plans that the search finds. At deployment, the learned worker replaces the search. We evaluate Landmarker on the IPC 2023 learning-track benchmark against state-of-the-art learned heuristics, learned policies, and LAMA. In the experiments, Landmarker solves 648 of the 900 test instances, more than every baseline except the strongest lookahead policy (668). Earlier landmark-based decompositions selected subgoals through search or heuristic rules. Landmarker makes landmarks viable as subgoals again by learning how to order and achieve them, solving more instances than LAMA, which uses landmarks only as a heuristic.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.