acceptodds
Under review as a conference paper at ICLR 2027

Where to Split? Benchmarking Subtask Splitting in Open-World Robot Data

Abstract

Subtask splitting is crucial for structuring long robot demonstrations into semantically meaningful units for robot learning. However, existing benchmarks and methods largely focus on constrained settings, limiting their applicability to rapidly growing and increasingly diverse open-world robot data. We introduce RoboSplitBench, a benchmark for systematically evaluating subtask splitting in open-world robot data. It contains demonstrations from diverse open-world scenes, with proprioception, multi-view observations, and human-annotated temporal subtask transition points. To establish a strong reference solution on this benchmark, we further propose RoboSplit. It is a training-free approach that explicitly assigns complementary roles to proprioception and Vision-Language Model (VLM) reasoning. Proprioception guides temporal windowing and selective use of different views, while VLMs predict subtask transition points. We systematically evaluate different methods on RoboSplitBench across diverse scenes and transition types, revealing the challenges of open-world subtask splitting. The strongest baseline achieves an F1 score of 0.479, while RoboSplit establishes a strong reference at 0.713 and consistently outperforms representative baselines across diverse scenes. As robot datasets continue to grow rapidly, our work provides a foundation for their structuring. We provide an anonymous link to RoboSplitBench's data, and an example data sample in supplementary materials.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.