GROW: Grounding with Long-Term Awareness for Skill-Incremental Robotic Manipulation
Abstract
Skill-incremental robotic manipulation requires continuously acquiring new skills while preserving previously learned capabilities. A widely adopted paradigm formulates manipulation as keyframe-wise action prediction, where the policy predicts the next keyframe action from the current observation and instruction. Applying skill-incremental learning to such keyframe-based manipulation introduces a fundamental level mismatch: local keyframe supervision and replay provide limited explicit guidance for preserving the long-term semantics of complete skill execution. To address this discrepancy, we propose **GROW** (**Gr**ounding with L**o**ng-Term A**w**areness), a continual manipulation framework that incorporates long-horizon action semantics across perception, action prediction, and replay, enabling skill-level semantics to be better captured and preserved across stages. At the perception level, GROW introduces dual-horizon grounding to regularize cross-modal attention toward both immediate interaction targets and long-horizon task goals. An auxiliary action chain predictor further captures how each skill evolves toward completion through instruction-conditioned future action prediction. The resulting action-chain predictions are additionally exploited for chain-based replay herding, enabling replay selection based on trajectory-level skill semantics. Experiments on RLBench simulation and real-world manipulation tasks show that GROW consistently improves performance over prior skill-incremental learning approaches. Code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.