Auditing a Learned Admission Gate for Offline-to-Online Safe RL: Distinct Certificates and Public Comparator Failures
Abstract
We audit one learned admission gate between BC-All proposals and BC-Safe fallbacks in offline-to-online safe reinforcement learning. We distinguish two certification targets. Fixed-time split-conformal accounting controls an immediate-cost proxy under deployment exchangeability and separately assumed fallback control. A weighted order-statistic rule instead bounds the marginal next-trajectory event \(\Pr(G_T>D,\mathcal Q\ne\varnothing)\le\alpha\) for a prespecified family of complete policies; it is not the high-confidence population-risk target of Learn then Test, and neither result protects qualification rollouts or repairs policy-induced shift. A synthetic study shows the intended mechanism and its pre-certificate exposure. On three public Safety-Gymnasium Point tasks, 54 completion/rejection-qualified BC-All/BC-Safe checkpoint pairs fail the frozen joint completion–violation criterion against a pre-treatment, deployable rate-matched mixer. A separate retrospective stress test permutes each gate trajectory's realized fallback mask at equal count and horizon; it is post-treatment and not a causal timing comparison. Gate violation is higher at all task-level point estimates under both comparisons, while completion often increases. Finally, an outcome-inspected diagnostic applying the trajectory rule to 600 retained gate trajectories per selected checkpoint yields an empty qualified set (0/54; rank 600; order statistics 134–810 for budget 10), under an unverified cross-campaign exchangeability premise. Thus the instantiated gate refuses trajectory certification and shows no joint public benefit; this does not generalize to a gate class or establish deployment safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.