When Is a Subtask Complete? Consistency-Aware Subtask Annotation for Hierarchical Robot Policies
Abstract
Hierarchical robot policies decompose long-horizon manipulation into a sequence of subtask to conduct a difficult task by high-level policy plans subtask instructions that a low-level policy executes. Training such policies requires decomposing long-horizon demonstrations into subtask segments, each labeled with an instruction and terminal state where the subtask is complete. In prior works, automated decomposition methods cannot ensure that the same subtask ends at the same state across demonstrations; one annotation may end "place the cup" at object release, while another ends it after arm retraction. Without a shared convention, the policy receives conflicting supervision about when to procede to the next subtask. In our pilot study, mixing two individually valid annotation sources lowers hierarchical policy success below that of either source alone (41.4% vs. 49.5% and 50.8%). To address this, we propose Consistency-Aware Subtask Decomposition (CASD), which learns an annotation convention from a small set of reference annotations and applies learned convention across the dataset. CASD over-proposes candidate boundaries from motion cues and scores every interval for the two properties a convention must fix: (1) which skill it shows and (2) whether it covers one complete occurrence of that skill. Because this completion criterion is learned from the references rather than inferred from per-episode visual change, it applies uniformly across demonstrations; a structured loss and dynamic programming then select a consistent labeled partition of each episode. On RoboCasa365 composite tasks, hierarchical policies trained on CASD annotations improve success rate by 10.0%p over the strongest automated annotation baseline and nearly matching simulator ground-truth annotations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.