Active Hypothesis Analysis for Language Model Reasoning under Ambiguity
Abstract
Reasoning from text about relationships among variables is challenging when multiple structural interpretations are compatible with the evidence. Given correlations and conditional independencies, a reasoning model must determine whether a proposed relation is entailed without knowing which interpretation is correct. We introduce Active Hypothesis Analysis for Language Models under Ambiguity (AHA), an inference-time procedure that explicitly represents competing interpretations and actively tests them. AHA combines a frozen language model with a deterministic controller: the model proposes candidate interpretations and answers targeted local questions, while the controller selects questions by expected information gain and maintains a posterior over hypotheses. We provide a kernel-based result relating representation similarity to differences in hypothesis scores, and a finite-budget guarantee for posterior concentration under candidate coverage, informative queries, and calibrated responses. On Corr2Cause, AHA achieves F1, compared with 51.6 for budget-matched self-consistency and 74.3 for random-query Bayesian aggregation using the same GLM-5.1 backbone. Local-question accuracy ranges from 62.4% to 85.9%, with concentration degrading as response noise approaches random guessing. We further introduce Extended Corr2Cause, an all-negative rejection stress test. Results on both benchmarks show that explicit uncertainty can direct inference-time computation toward questions that distinguish competing interpretations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.