Reliable AI-Assisted Decisions via Selective Auditing
Abstract
We consider online monitoring of an AI system that produces a sequence of decisions. At each round, the system outputs a proposed decision together with a confidence score estimating the likelihood that the decision is trustworthy. The decision and confidence score need not be produced by the same model: for example, one agent may propose a decision, while a separate verifier agent evaluates it and outputs a correctness or safety score. Based on this information, we must decide whether to accept the AI decision automatically or consult an expert, such as a human, who reviews the proposal and either accepts or corrects it. We formulate this task as an online constrained optimization problem in which the goal is to minimize expert auditing while guaranteeing that the frequency of unaudited AI errors remains bounded. We derive a family of algorithms with provable guarantees on unaudited errors, including a method that is optimally efficient: it achieves the theoretical minimum auditing rate for both adversarial and i.i.d. data streams. Experiments on detecting harmful requests to large language models and on qualitative text annotation show that our methods achieve valid error control in practice while substantially reducing the need for human auditing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.