Learning to Test: Sequential Monitoring of Generative Models and Agents
Abstract
As AI models scale and are deployed with increasing autonomy, their opaque decision-making can pose risks to users and providers, motivating guardrails and continuous task monitoring for compliance with safety and deployment protocols. However, monitoring quality can fluctuate depending on the model and problem at hand, and generally lacks formal assurances. We recast deployment monitoring as a statistical testing problem, viewed in particular through the lens of sequential hypothesis testing, or 'testing by betting'. Leveraging the framework’s error-control qualities, we equip monitors with verifiable false alarm controls such that benign model behavior triggers interventions only at a prespecified, tolerable rate. Challenged by the ambiguity of defining complex black-box monitoring problems in statistical terms, we investigate data-driven test designs with learned reference, betting, or evidence-normalization components. Empirical evaluations span diverse models and monitoring targets, including language generation, chain-of-thought reasoning, multi-turn interactions, and agentic tool use. Our results affirm that learned tests exhibit error-controlling and adaptive properties, while highlighting the nuances of translating raw monitoring scores into statistically valid evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.