True Claims, False Answer: Spot-Check Debate Is Limited by What the Honest Side Claims
Abstract
Spot-check debate lets a weak judge verify only the values that debaters choose to claim, on the premise that a false answer must rest on a checkable false claim. A dishonest debater that claims only values shared by the true and false computations defeats this premise: every check confirms it. Against such an attacker, a check helps only in a covered debate, where the honest debater claims a discriminating value (one on which the two computations differ) and the judge checks it. By construction no check contradicts the attacker, which makes it an instrument for measuring the honest side and the judge. We build a testbed that records, for every debate, whether the honest side held such a value, claimed it and had it checked: 600-step arithmetic programs that an oracle verifies step by step but the judge cannot execute, with one LLM in all roles. With the full true trace, one check suffices (accuracy 0.98). With a thinned trace, coverage is lost on the honest side: our open-weight honest debaters claim a discriminating value in only a third to half of the debates in which they hold one. The judge checks mostly what it is asked to check and, without coverage, sides with whichever debater a check confirms, a rule that is Bayes-optimal against ordinary liars. Against ordinary liars the failure does not show, because their own false claims get checked: when the honest side holds only every 60th value, the same honest debater and judge score 0.86 against an ordinary liar but 0.53 against the attacker, close to the 0.47 of no debate. The gap replicates with Qwen2.5-72B in every role. With every 10th value, scripting the honest side to claim all it holds and request checks of its latest values nearly doubles coverage and raises accuracy from 0.60 to 0.84, only part of that gain through coverage. Evaluations of debate should include such an attacker and report how often the honest side claims the discriminating values it holds and how often these are checked.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.