PACE: Discovering Options via Self-Distillation
Abstract
Canonical reinforcement learning (RL) policies require reading an observation before sampling an action. However, running a policy at each step for a single decision can be expensive. Should an agent be able to emit n actions at a time, this would allow reducing the number of observations it has to read and the number of inference passes it has to run during a decision process, unlocking large computational savings. In this work, we argue that many decision problems contain states where one observation carries enough information for the policy to emit more than one optimal action at a time, making intermediate state reads an avoidable cost. We propose PACE (Per-state Adaptive Commitment): while Proximal Policy Optimisation (PPO) trains as usual, PACE distills it into a multi-step policy that reads one observation and emits several consecutive actions from it alone. Every time this policy reads an observation from the environment, it also predicts the number of actions to output by measuring how far that replay can be trusted, and committing exactly that far. On random tasks with deterministic transitions, PACE reduces the number of required observations by up to 8x, and the compression shrinks as transition stochasticity rises. Furthermore, we show that, in open-ended settings, recursively augmenting the action space of a Markov Decision Process with the options discovered by PACE drastically improves success rates in environments where both curriculum and tabula-rasa learning fail.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.