acceptodds
Under review as a conference paper at ICLR 2027

Every Failure Is a Timeout: Failure-Mode Homogeneity Makes Robot Policy Evaluators Unverifiable

Abstract

Automatic robot policy evaluators are commonly validated by their agreement with ground-truth success labels. We show that such agreement can provide no evidence of genuine evaluation capability when the evaluation protocol itself reveals the outcome. Under stop-on-success evaluation, policies without a termination channel fail exclusively by reaching the time horizon, allowing episode duration alone to distinguish success from failure. We establish this failure mode through a controlled intervention on the termination channel and construct a frame-blind witness that exactly reproduces a vision–language judge's evaluation metrics. We further formalize an evidence budget based on success–failure pairs misordered by duration and derive a randomization floor below which evaluator certification is impossible. Across 47 screened conditions, only two satisfy this floor, and neither remains admissible after suppressing the termination channel. To address this identifiability failure, we investigate fixed-horizon logging, which eliminates duration-based outcome leakage without increasing the number of rollouts. This protocol also distinguishes achievement from terminal success, revealing discrepancies concealed by stop-on-success evaluation. However, removing duration leakage does not guarantee evaluator verifiability: the first success time retains predictive information, while some terminal success labels depend on historical states inaccessible to frame-based evaluators. Together, our findings establish that reliable evaluator validation requires both sufficient evidence beyond protocol-induced shortcuts and alignment between ground-truth labels and observable information.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.