TraceCut: Atomic Action Segmentation and Labeling for Robot Manipulation Videos via Actionability-Aware Hierarchy Decoding
Abstract
Long-horizon tasks and their corresponding planning models represent a prominent yet challenging frontier in robotics research. While collecting diverse long-horizon task videos has become a key focus, fine-grained action segmentation and annotation remain fundamental to constructing high-quality datasets. Although temporal grounding based on vision-language models (VLMs) offers a natural solution, it often struggles inferior accuracy due to the burden of simultaneously handling action perception, temporal localization, and semantic labeling. To address this, we propose TraceCut, a video atomic action segmentation and annotation pipeline based on hierarchical segment decomposition. TraceCut builds a hierarchical action tree for robot videos to enable accurate multi-level segmentation. It achieves precise atomic action splitting and annotation by generating and aggregating segment-level descriptions. Furthermore, we develop a lightweight video captioning model to fulfill the critical requirement for efficient and accurate caption generation within the pipeline. We also introduce RoboActionTrace, a standardized benchmark for evaluating robot video action segmentation and annotation. Experiments demonstrate that both the TraceCut pipeline and the lightweight captioning model achieve competitive performance in atomic action segmentation and annotation. Moreover, data constructed via this pipeline significantly enhances performance in core long-horizon task planning embodied task, demonstrating substantial commercial potential amid the growing demand for automated annotation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.