Robust Information-Gain Control Active Agentic Reasoning under Approximate Beliefs
Abstract
Large language model (LLM) agents in multi-turn reasoning often fail to make epistemic progress under approximate belief tracking. We formalize this failure through Belief Trap Regions (BTRs), forward-invariant regions of belief space in which expected progress stagnates. We show that even information-gain maximization (IG-max) can induce such traps under belief approximation error: small perturbations in the agent’s internal belief may invert action rankings, causing repeated selection of uninformative actions despite the existence of informative alternatives. To mitigate this failure mode, we propose Distributionally Robust Information Gain (DR-IG), a decision-time objective that optimizes worst-case information gain over an ambiguity set around the current belief estimate. We show theoretically that robust worst-case control restores the negative drift of a progress potential under bounded belief error whenever uniformly informative actions exist. We further introduce a lightweight verbalized belief proxy that extracts approximate belief states from confidence-annotated model outputs, enabling belief-aware planning without additional training. Across interactive reasoning benchmarks and multiple LLM families, naive IG-based agents frequently collapse into deterministic self-locking loops with near-zero solve rates. In contrast, DR-IG substantially reduces trap incidence and partially restores epistemic progress under approximate beliefs. Together, our results identify a concrete mechanism underlying active-reasoning failure in LLM agents and demonstrate that robust belief-aware control can mitigate these pathologies in a training-free setting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.