acceptodds
Under review as a conference paper at ICLR 2027

Planning Guided Option Critic

Abstract

A symbolic skill specification describes when a skill may run, what it must achieve, and how it is rewarded, but says nothing about how to learn the skills or how to choose among them. We introduce Planning Guided Option Critic (PGOC), which learns the skills and an option-critic controller jointly from the specifications alone, deriving masks and plan-level rewards directly from them. Preconditions and continued-need predicates restrict which options may be selected; plan-progress rewards credit the controller for satisfying remaining requirements; and each completed skill bootstraps from the controller's continuation value, crediting a skill for the opportunities it creates rather than its position in a fixed plan. Across five training seeds on Craftax, PGOC with recurrent backbones reaches the Gnomish Mines in 62.8–64.5% of frozen evaluation episodes, versus 7.7–9.3% for SCALAR and at most 4.3% for PPO. Ablations that retain the critic while varying the selection rule decompose this gain by task structure. On Craftax, where a fixed plan already suffices, prescribing that plan raises Mines success further to 69.4%: the credit assignment, not the learned ordering, drives the fixed-plan task. On randomized TreasureDash, where the optimal plan depends on the encountered state, the learned controller is decisive, scoring 22.2 against 10.9 for the prescribed order and 13.3 for uniform selection. The same option critic thus supplies a single credit assignment for both kinds of task, while the value of learning the selection rule itself depends on the task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.