acceptodds
Under review as a conference paper at ICLR 2027

Learning State-Dependent Reach-Avoid Bounds with PAC Guarantees

Abstract

Learned control policies can succeed on average yet fail from particular initial states. Existing reach-avoid certificates either require a model of the dynamics with Lipschitz bounds and a grid over the state space, or certify a single success probability. We introduce Reach-Avoid Probability Certification (RAPC), which learns a function that lower-bounds the probability that a policy started at reaches a target before an unsafe set, and certifies it from simulator rollouts alone. With probability at least over learning and certification, a returned overstates the true success probability only on initial states of total probability at most , even when the policy and bound are updated adaptively. The key step links an unobservable event, an overstated bound at a start, to an observable one: a failed binomial confidence check on repeated rollouts from that start. Calibrating to these confidence limits gives a learner-independent lower bound on its acceptance probability and, for a fixed calibration pipeline, certifies without a separate validation batch. On a 12-dimensional Hopper environment and a PointGoal environment with 44-dimensional observations, RAPC certifies bounds in all ten seeds of each task. On a wider Hopper start distribution, the learned bound averages over ten seeds, against for the largest constant accepted by the same test. We release our code at https://anonymous.4open.science/r/iclr-pac-BBBD/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.