Controlling the Probability of Unknown Unknowns in Pre-Deployment Testing of AI Systems
Abstract
In pre-deployment testing of AI systems such as AI agents, a developer can often deal with *seen* failure modes — leaking API keys, deleting important databases, or installing a malicious package — by implementing fixes or making contingency plans for when the same issue might occur again. What is more challenging are the *unknown unknowns:* the *unseen* failure modes that haven't surfaced during testing. The decision to stop testing and deploy therefore faces an inherent tradeoff: continue testing to observe more behaviors, at the cost of time and resources, or stop testing and deploy, at the risk of failure modes remaining undiscovered. How then does the developer decide when they have tested enough? To answer this, we develop an algorithm that certifies when the *probability of an unseen failure mode* is sufficiently small. The algorithm constructs an anytime-valid upper confidence sequence on this quantity, so the developer can repeatedly test a system, inspect the bound, and decide to stop when it falls below a chosen threshold without invalidating the statistical guarantee. Our problem can be interpreted as a sequential variant of the missing-mass estimation problem which gave rise to the classic Good-Turing estimator (Good, 1953). On SWE-Chat, a dataset of real-world coding agent traces (Baumann et al., 2026), we implement our procedure and show that it achieves tighter upper bounds than a simple multiple-testing corrected Good-Turing baseline, and even dominates a fixed-horizon (non-multiple-testing-corrected) Good-Turing upper confidence bound, which is not anytime valid. Additionally, to better match realistic evaluation pipelines, we extend our method to handle sampling from a different task distribution and imperfect automated failure labeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.