When the Opposing Agent Knows: Evidence Disclosure in LLM Debate
Abstract
LLM debaters may leave out evidence that weakens their assigned positions. We ask whether a debater is more likely to introduce such evidence first when it is told that its opponent also holds it. We study this question in DisclosureBench, which covers financial investment, football prediction, real-estate valuation, and cardiac assessment. In each instance, two debaters argue opposing positions, and each evidence item is randomly given to both debaters or to only one of them; predefined labels mark which items are unfavorable to each position. From the chronological dialogue, we measure how often a debater introduces its own unfavorable evidence before its opponent does, and compare items shared with the opponent against items it holds alone. In all four domains and for both roles, debaters introduce shared unfavorable evidence first far more often than private evidence (by 16 to 60 percentage points). The gap holds at every sharing level and allocation seed in all four domains, and across three models in matched financial comparisons. In long financial debates, the gap appears only after a debater has seen at least one opponent turn. Debaters thus bring up unfavorable evidence mainly when the opponent could raise it anyway, so the evidence most likely to go unsaid is the evidence only one debater holds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.