ROBOLEAK: CLAIM-CONDITIONED OVERLAP AUDITING FOR EMBODIED AI BENCHMARKS
Abstract
Robot-learning benchmarks often reuse objects, scenes, and skills across training and evaluation to probe transfer under controlled changes. Yet a benchmark score alone does not reveal which aspects of the evaluated behavior are novel relative to the training collection. We introduce ROBOLEAK, a claim-conditioned auditing framework that first records cross-boundary relations at five levels—stored files, decoded observations, near-duplicate views, object and scene structure, and task semantics—and then applies a declared evaluation contract specifying which observed relations violate the benchmark’s intended novelty. Auditing LIBERO with ROBOLEAK shows that all seven Goal-task object models already appear in the training suites, while no scene composition is repeated; under a strict novelty rule a single familiar task is excluded, and this one exclusion raises SmolVLA’s score from 76% to 80%, leaves X-VLA unchanged at 100%, and slightly lowers VLA-JEPA’s score. The same cleaning moves different models in opposite directions because it removes a task each model finds differently hard. A complementary controlled exposure study finds no consistent average gain from exact-task demon- strations, while revealing substantial seed-level variation under a fixed training budget. ROBOLEAK turns a benchmark score into an answerable question: what was shared, what was required to be new, and which tasks does the number average over?
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.