acceptodds
Under review as a conference paper at ICLR 2027

Factored Q-Chunking

Abstract

Reinforcement learning (RL) from demonstrations can teach robots long-horizon manipulation skills despite sparse rewards. Q-chunking, a simple and effective recipe for this setting, has the policy output a short sequence (a chunk) of future actions and scores the whole chunk with one critic. However, it commits every joint of the robot to the same chunk, even though many robots combine parts that act on different time scales, such as a mobile base carrying two arms, or an arm carrying a dexterous hand. Letting some joints re-plan mid-chunk breaks value learning, because a critic over the full chunk then scores actions that are never executed. We propose Factored Q-Chunking (FQC), which commits a slow part of the action for the full chunk and re-plans a fast part halfway through. FQC gives each part its own critic that scores exactly the actions that part executes. On DexJoCo (arm and hand) and BiGym (base and two arms), FQC substantially improves online fine-tuning over Q-chunking under the same policy and budget. Ablations show that matching each critic to what is actually executed is essential.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.