Information-Limited Black-Box Safety Evaluation under Latent Deployment Context
Abstract
Black-box safety evaluation observes model behavior under an evaluation distribution but aims to estimate risk under deployment, where consequential triggers may be substantially more prevalent. We study a common-support hidden-set construction with trigger prevalence during evaluation and during deployment. Under identical non-trigger observations, an adaptive evaluator cannot localize the hidden trigger set through pre-discovery interaction alone; systematic improvement over the zero-information baseline requires informative side information. For distinct queries, we prove that the true trigger-discovery probability and its independent-side-information baseline satisfy , where denotes binary KL divergence and quantifies information about the hidden trigger set. A counterfactual-query coupling reduces adaptive discovery to a binary event and yields an information-dependent minimax lower bound for deployment-risk estimation. We further establish an approximate-symmetry extension and characterize the discovery–information trade-off for an analytically tractable hint channel, including its weak- and strong-hint limits. Controlled simulations, a 140-configuration parameter sweep, and access-hierarchy experiments illustrate the finite-sample implications of the theoretical construction. A preliminary trained-model diagnostic illustrates why coarse approximate-symmetry assumptions can yield vacuous bounds at practical query budgets. Our contribution is primarily theoretical: it makes evaluator access assumptions explicit and separates query budget from information about deployment-relevant structure. End-to-end validation on language models, including the proposed open-weight protocol, remains an important direction for future work.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.