acceptodds
Under review as a conference paper at ICLR 2027

Efficient Scalable Oversight via Rewinding

Abstract

Scalable oversight asks how a weak verifier can check the outputs of computations performed by more powerful AI systems. One proposed approach uses relativizing interactive proofs, in which a weak verifier checks the output of a complex oracle-aided computation by interacting with one or more powerful provers. A prominent example is debate, which is highly efficient but assumes that one of two powerful AI models, or provers, is truthful. For certain classes of oracle-aided computations, single-prover relativizing proofs remove this assumption, but at substantially greater computational cost. In this work, we present a new approach to scalable oversight based on relativizing multi-prover interactive proofs (MIPs). We construct a relativizing MIP for computations whose outputs are robust to changing a small fraction of oracle answers. Our construction improves the efficiency of prior single-prover relativizing proofs without assuming that any prover is truthful. Instead, it uses two provers that may coordinate beforehand but cannot communicate during verification. This separation can be realized by interacting with isolated copies of an AI model, or by effectively “rewinding" a single model, provided its state can be reset so that each interaction has access only to its prescribed history. In our MIP construction the honest provers run in nearly the time of the original computation, and the verifier (after a one-time preprocessing step) runs in near-linear time in the input length. We implement our MIP and report results on a preliminary MNIST experiment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.