Beyond Repeatable Scores: Experimental Fidelity Auditing in Robot Benchmarks
Abstract
A repeatable benchmark score is meaningful when its scenes, attempts, and observations preserve the declared experiment. Yet aggregation can conceal changes to these units, masking a different evaluation. We study this gap through experimental-unit fidelity and trace how execution and reporting can alter the declared units. This perspective leads to EFA-Replay, which aligns records with declarations, tests evaluator changes with matched controls, and determines which results the evidence supports. Paired-label accounting exposes changes hidden by aggregation, while sharp compatible-record bounds identify what reduced records preserve. In an ACT evaluation, scene count and 80% success remain unchanged despite 20% request-label disagreement. For SimplerEnv, exact pooled success leaves condition variance unidentified in 64/78 cells. A separate LIBERO audit flags a stale first observation before scoring; correcting it changes ACT task-8 success from 11/30 to 27/30 on frozen follow-up states. These results motivate preserving the unit-level evidence that gives benchmark scores their meaning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.