acceptodds
Under review as a conference paper at ICLR 2027

Equilibrium of Thoughts: Reasoning as Iterative Stabilization Under Adversarial Probing

Abstract

Reasoning frameworks from Chain-of-Thought to Tree-of-Thoughts are constructive: they build an answer step by step, so a single early error cascades and is rarely caught. These methods find an answer but cannot certify it. We propose Equilibrium of Thoughts (EoT), which reframes reasoning as stabilization under adversarial probing: generate an answer, attack it for the single strongest objection, revise if the attack lands, and repeat until no valid objection remains. The surviving answer is the reasoning equilibrium, grounded in fixed-point, argumentation, and game theory. Our central finding concerns who should attack. A model cannot find the flaws it just made, so self-critique is unreliable. The bottleneck is the critic. Holding the reasoner fixed and varying only the adversary, we show a strong, decorrelated cross-model critic is the active ingredient: it rescues a weak reasoner ( on AIME for a 32B model) and lets critic capability substitute for reasoner capability, while a same-strength same-family critic does not, which isolates the contribution of decorrelation from that of raw strength. Recast as cross-model adversarial verification, this prompt-only recipe helps across math, science, code, and multi-hop QA, including benchmarks where multi-agent debate degrades accuracy. Finally, EoT wraps an initializer instead of replacing it: composed on Tree-of-Thoughts with the same critic, it beats debate on every hard benchmark (AIME-2024 vs. ). Adversarial verification thus complements constructive reasoning as a primitive that amplifies it: the stronger the initializer, the higher EoT reaches.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.