acceptodds
Under review as a conference paper at ICLR 2027

SCHEMA: Scalable Harness-Guided Temporal Quality Metadata Construction for Robotic Manipulation Learning

Abstract

Robotic manipulation trajectories contain temporally heterogeneous behavior, including productive progress, inefficient attempts, recoveries, and failures, yet standard Vision-Language-Action (VLA) model training largely treats these segments alike. We introduce **SCHEMA**, a scalable harness-guided framework that uses a Vision-Language Model (VLM) to construct sparse temporal quality metadata through iterative evidence seeking, contextual quality assessment, and boundary refinement, enabling better use of trajectory quality during VLA training. We further introduce **RTQ-Bench** (Robot Temporal Quality Benchmark), built from multiple public manipulation datasets with human-annotated quality intervals. On RTQ-Bench, SCHEMA substantially improves temporal quality localization over single-pass inference with the same backbone and performs competitively against diverse VLM- and score-based baselines. Pretraining on approximately 1,500 hours of real-world manipulation data with SCHEMA-guided sampling improves downstream VLA performance over uniform sampling, with the training data, model initialization, action-learning objective, and optimization schedule held fixed. Across four real-world tasks, SCHEMA-generated metadata improves both sampling-based and advantage-conditioned learning, with larger gains when failed policy rollouts are included.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.