Task-Definition Conflicts in Robot Benchmarks: Measurement and Learning
Abstract
Robot benchmarks distribute task definitions across instructions, demonstrations, rewards, and success checks. These components can agree at runtime yet implement a goal different from the declaration. We introduce executable task contracts to trace this correspondence and separate measurement changes from learning interventions. The same reference guides Contract-Guided Supervision Repair (CGSR), which searches late-to-early rollback boundaries, preserves action prefixes, and verifies complete repaired executions. In LIBERO, native and calibrated completion checks differ by 38.0 percentage points on identical executions. In a frozen extension, CGSR repairs nine of nine fresh-execution failures versus six for full recollection, while halving candidate-execution cost per verified repair. A subsequent exploratory student comparison improves calibrated completion by 13.3 points under equal training updates and unequal acquired-data yields. Blinded human review and RoboCasa monitor diagnostics clarify the interpretation of these outcomes. Together, explicit task-reference tracing and verified repair connect benchmark claims to the behaviors actually measured and taught.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.