Can AI Oversight Be Zero-Knowledge?
Abstract
AI systems increasingly produce outputs from confidential data: a fitness-for-duty assessment from an employee's medical records, or the predicted properties of a drug candidate from its molecular structure, a trade secret. It is important to verify that such outputs are correct without revealing the data on which they are based. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or an external database such as the web. These works focus on efficient verification, which was shown to be impossible for general oracle-aided computation and so requires additional assumptions on the model. In this work, we focus instead on privacy. To sidestep the impossibility, we allow the verifier to run in time polynomial in the computation, and we ask whether interactive arguments for oracle-aided computation can be zero-knowledge: the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there is no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge, assuming only collision-resistant hash functions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.