acceptodds
Under review as a conference paper at ICLR 2027

Cryptographic Techniques for AI Safety via Debate

Abstract

AI debate aims to let a weak judge evaluate stronger AI systems by having competing agents expose each other’s errors. We study two mechanisms for strengthening debate: cryptographic commitments and private communication. Building on classical competing-prover protocols, we give debate protocols whose communication and judging costs scale polylogarithmically with computation length, in addition to input and security-parameter costs. These guarantees extend the standard debate model from PSPACE to EXP and accommodate black-box calls to tools, databases, or human feedback. With cryptographic commitments and random access to the input and transcript, the judge checks only a constant number of locations and makes at most one black-box call. We evaluate private communication in LLM debates where the agent defending an incorrect answer is stronger than its opponent and the judge; experiments show that privacy in debate can increase the protocol's accuracy and its robustness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.