Factored Q-Chunking
Abstract
Reinforcement learning (RL) from demonstrations can teach robots long-horizon manipulation skills despite sparse rewards. Q-chunking, a simple and effective recipe for this setting, has the policy output a short sequence (a chunk) of future actions and scores the whole chunk with one critic. However, it commits every joint of the robot to the same chunk, even though many robots combine parts that act on different time scales, such as a mobile base carrying two arms, or an arm carrying a dexterous hand. Letting some joints re-plan mid-chunk breaks value learning, because a critic over the full chunk then scores actions that are never executed. We propose Factored Q-Chunking (FQC), which commits a slow part of the action for the full chunk and re-plans a fast part halfway through. FQC gives each part its own critic that scores exactly the actions that part executes. On DexJoCo (arm and hand) and BiGym (base and two arms), FQC substantially improves online fine-tuning over Q-chunking under the same policy and budget. Ablations show that matching each critic to what is actually executed is essential.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.