Chunking the Policy: Extending the Policy-Update Horizon with a Compact Transformer
Abstract
We introduce HOPS (Horizon-Extended Optimization over Policy-Generated Sequences), an off-policy actor–critic method that independently controls the policy-update horizon while keeping interaction-time prediction and execution horizons fixed. A compact Transformer autoregressively composes action chunks of length , a task-dependent hyperparameter, into longer sequences for actor optimization while retaining short prediction and execution horizons during interaction. In the default Meta-World configuration, : actor updates optimize sequences of up to sixteen actions, while interaction predicts and executes (default: 4) actions before replanning. Twin prefix-conditioned critics provide value gradients across the generated sequence. We train separate agents online from scratch on all 50 Meta-World ML1 tasks and dense-reward FANCY GYM Box Pushing, using final-timestep success for evaluation. HOPS achieves a mean task-wise success IQM of on Meta-World at environment interactions, compared with for T-SAC, the previous State-of-the-art baseline. On Box Pushing, HOPS reaches a success IQM of 0.941 at interactions, compared with 0.637 for T-SAC. Controlled ablations on a 20-task Meta-World subset vary the maximum policy-update horizon while holding the interaction horizons and critic training fixed. Extending the maximum update horizon from four to sixteen improves aggregate performance, with task-dependent effects. These results support extending the temporal scope of actor optimization through chunk-autoregressive generation while retaining short action chunks for interaction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.