Learning Chunk Values from Partial Execution in Offline-to-Online Reinforcement Learning
Abstract
Action chunking, in which a policy generates sequences of actions rather than individual actions, can improve the sample efficiency of reinforcement learning through temporally coherent exploration and multi-step value learning. Adaptive action chunking selects how many actions to execute based on the current state, allowing the agent to replan using new observations. Though only part of a full chunk is executed, the policy must learn to generate a complete action chunk. Learning the value of that chunk from partial execution can provide feedback for improving the entire sequence. In this paper, we propose Partial-Execution Q-Chunking (PEQC), an offline-to-online reinforcement learning method that learns the values of complete action chunks from partial execution. PEQC retains generated chunks and constructs additional learning targets for the full-chunk value used in policy optimization by combining observed rewards with the estimated value of the original unexecuted actions. On OGBench, PEQC improves average evaluation success during online fine-tuning on several manipulation and navigation tasks relative to the same agent without this additional supervision. The observed gains suggest that full-chunk values can be effectively learned from partial execution while retaining the flexibility to replan before completing the generated action sequence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.