acceptodds
Under review as a conference paper at ICLR 2027

Learning Heuristic Functions via Policy Optimization for Symbolic Planning

Abstract

Tree search efficiency in deterministic symbolic planning critically depends on heuristic functions, which are difficult to design. Heuristics are therefore often learned, and one promising way to train them is with ranking losses, which teach the heuristic to rank the states in the open list, on problem instances whose solutions are known. We consider the opposite scenario, in which no solutions are available and the heuristic is trained on the searches it performs itself. The training data then depends on the heuristic, which turns learning into a policy optimization problem. We revisit heuristic learning from a policy optimization perspective and formulate symbolic planning search as a Markov Decision Process (Tree-MDP), where a state is a partially expanded search trees and an action selects a node for expansion. The Tree-MDP allows us to (i) optimize heuristics with standard reinforcement learning algorithms, which lets us compare the Tree-MDP with the standard formulation of planning as an MDP using the same algorithm, (ii) show that ranking losses are closely related to policy optimization in the Tree-MDP, and (iii) train the heuristic on the search configurations it visits itself rather than only on those along the best path found. Experiments on IPC 2023 learning track benchmarks support the theoretical connections : PPO performs better in the Tree-MDP than in the standard formulation, ranking losses and their policy counterpart perform alike, and training on the heuristic’s own search configurations yields the highest average coverage among the learned heuristics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.