acceptodds
Under review as a conference paper at ICLR 2027

Rollout-Level Calibration for Failure Monitoring in Generative Robot Policies

Abstract

Failure monitors for generative robot policies must raise timely alarms while limiting false positives on executions that would otherwise succeed. We introduce RoLCA (Rollout-Level Conformal Alerting), a modular layer that aggregates progress-matched score evidence over time and across channels, then split-conformally calibrates its maximum over eligible monitoring times using successful rollouts only. Under exchangeability of successful rollouts, RoLCA provides finite-sample rollout-level false-positive control for causal scoring, aggregation, fusion, and monitoring rules fixed independently of threshold calibration. The guarantee is marginal over the calibration draw and the next successful rollout, without requiring within-rollout independence. Across 30,502 rollouts in six task–backbone cells, calibrated monitors operate near the nominal false-positive rate in two sealed confirmations, but detection gains depend on the monitoring convention. On the second confirmation, a post-hoc comparison yields a 7.6-percentage-point paired recall gain over matched running-average OR under fixed-horizon monitoring; this advantage is not established after stop-at-success recalibration. Fusion improves early recall over a rollout-trained single channel but does not consistently outperform execution consistency alone. A clock-only baseline detects all timeout failures in this population under stop-at-success evaluation, motivating assessment of early warning beyond eventual coverage. These results separate false-alarm validity from detection power and underscore the need for explicit monitoring intervals, matched comparators, and timing baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.