Compiling Task-Event Circuits for Offline Goal-Conditioned Reinforcement Learning
Abstract
Goal-conditioned policies often reach nearby goals yet fail when a task requires composing many distinct changes. We present the Algebraic Task Quotient (ATQ), a compiler that extracts repeated task changes from offline replay. ATQ represents task progress with persistent variables, which we call registers. It learns how recurring task-level events update these registers and compiles the learned rules into a task-event circuit. Each path through the circuit represents a possible sequence of task changes. At test time, a planner selects an event sequence, a policy executes one local request, and feedback updates the circuit state. On all eight OGBench Puzzle tasks, ATQ reaches 95.3–100.0% success with eight independently trained GCIQL policies. Under the same policies and test conditions, Test-Time Graph Search reaches 0.2–44.1% on the 4x4–4x6 layouts, whereas ATQ reaches 95.3–100.0%. ATQ also exceeds the cited published scores on all 4x5 and 4x6 tasks. On the larger puzzles, 98.4–100.0% of its requests apply an event in a register configuration where that event was absent from compiler training data. Episodes composed entirely of these new combinations still succeed 95.3–100.0% of the time. Beyond Puzzle, ATQ raises success in navigation, object manipulation, and soccer, with gains of up to 27.2 percentage points. These results show that replay can reveal event rules that support composition beyond search over recorded states.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.