SCHEMA: Scalable Harness-Guided Temporal Quality Metadata Construction for Robotic Manipulation Learning
Abstract
Robotic manipulation trajectories contain temporally heterogeneous behavior, including productive progress, inefficient attempts, recoveries, and failures, yet standard Vision-Language-Action (VLA) model training largely treats these segments alike. We introduce **SCHEMA**, a scalable harness-guided framework that uses a Vision-Language Model (VLM) to construct sparse temporal quality metadata through iterative evidence seeking, contextual quality assessment, and boundary refinement, enabling better use of trajectory quality during VLA training. We further introduce **RTQ-Bench** (Robot Temporal Quality Benchmark), built from multiple public manipulation datasets with human-annotated quality intervals. On RTQ-Bench, SCHEMA substantially improves temporal quality localization over single-pass inference with the same backbone and performs competitively against diverse VLM- and score-based baselines. Pretraining on approximately 1,500 hours of real-world manipulation data with SCHEMA-guided sampling improves downstream VLA performance over uniform sampling, with the training data, model initialization, action-learning objective, and optimization schedule held fixed. Across four real-world tasks, SCHEMA-generated metadata improves both sampling-based and advantage-conditioned learning, with larger gains when failed policy rollouts are included.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.