Disagreement Graph Inference
Abstract
Self-Consistency decoding improves large language model reasoning by sampling multiple Chain-of-Thought chains and selecting the most frequent final answer via majority vote, but treats each chain as an independent vote, discarding shared intermediate reasoning structure. We introduce Disagreement Graph Inference (DI), a training-free framework that replaces majority voting with structured probabilistic inference over intermediate sub-answer variables. DGI decomposes sampled chains into a sub-answer matrix, estimates pairwise dependencies via bias-corrected mutual information with false discovery rate control, constructs an optimal Chow-Liu tree, and performs exact belief propagation to compute posterior marginals. The key mechanism-the rescue effect—allows strong intermediate consensus to override an incorrect final-answer plurality through message passing along dependency edges. We prove formal rescue guarantees on trees, characterize algebraically symmetric conditions under which DGI can harm accuracy, and show that DI strictly generalizes Self-Consistency. Experiments across six reasoning benchmarks and three foundation models demonstrate accuracy improvements of up to 6.3% absolute over Self-Consistency with less than 150 ms overhead per instance, rescue rates up to 39.5%, calibration error reductions exceeding 50%, and graceful degradation on structurally trivial problems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.