acceptodds
Under review as a conference paper at ICLR 2027

What Zero Detections Certify: Statistical Safety Cases for Model-Aware Monitoring

Abstract

Monitoring is now one of the main safeguards for deployed AI systems. A monitor reviews only a subsample of traffic, and zero or few detections are often treated as evidence that harmful behavior is rare. That reasoning is not always valid: it assumes that the monitored system cannot use features of the context to infer where oversight is weak. We study monitoring-based safety cases when this assumption fails. Under invisible per-unit randomization and audited lower bounds on recall, we give finite-sample upper confidence bounds on the number of harmful actions. The bounds make no stochastic assumption on the system's behavior, but they depend on a sensitivity parameter that limits how much harmful behavior can concentrate in weakly monitored traffic. In a simple stratified design, a 95% bound of about three harmful actions under proportional allocation becomes 244 when harms may be twice over-represented in weakly monitored traffic. Using the same budget to equalize effective coverage instead gives a bound of about eight that is robust to . When the system only partially observes the monitoring policy, the relevant quantity is not average coverage, AUC, or guessing accuracy. It is the extreme lower tail of the system's posterior detection probability. We show that standard discrimination summaries cannot control this tail. Subsampled monitoring also cannot rule out a single catastrophic action unless per-unit detection is nearly complete, but it can bound the chance that a multi-step campaign escapes interruption. These results lead to a monitoring design principle: make knowledge of the policy strategically unhelpful, by equalizing effective coverage or adding a hidden randomized coverage floor. We turn the theory into a reporting protocol and illustrate it on public descriptions of monitoring systems from Anthropic and OpenAI. On three public attack benchmarks scored by seventeen open-weight monitor configurations, the recall of a single monitor ranges from to across attack strata, and the certificate is set by the strata where recall is lowest.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.