acceptodds
Under review as a conference paper at ICLR 2027

Acting in the Right Basis: Test-Time Compositional Generalisation for Offline Goal-Conditioned RL

Abstract

Robotic manipulation goals combine several factors, such as the placements of different objects. Offline data cover each factor but rarely their combination, so a goal-conditioned agent must generalise compositionally. Test-time methods improve a frozen agent, the backbone, without new data, yet they compose along time. They chain dataset states by graph search or fine-tune on retrieved sub-trajectories. When no dataset state realises the goal combination, the chain keeps a long final hop. Retrieval then teaches nothing beyond the backbone value. We compose along factors instead. Our method, GC-FACT, shows the frozen policy one single-factor target at a time. Such a target copies the current state and moves one object to its goal placement. Play data contain many such changes. This works only if the policy can carry out each target. That depends on how the goal is split into units, which we call the basis. We prove that the serial plan is near optimal when units are atomic and weakly coupled. We also prove that the relative cost of each target depends on the basis, not only on the environment. On the Lights Out puzzles, per-cell targets are unreachable or costly. Per-press targets are cheap, yet monotone descent over them stalls. A linear solve over the binary field finds the exact plan. GC-FACT lifts three frozen backbones on nine OGBench manipulation tasks. Their average success rises from 34.4, 9.3 and 5.0 percent to 60.7, 45.0 and 32.6 percent. Graph search lowers all three averages.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.