Investigation Capability versus Investigation Intelligence: Diagnosing Selective Evidence-Gathering in LLM Agents with Information-State Oracles
Abstract
LLM-controlled agents can gather costly diagnostic evidence before updating a belief; this capability alone is consequential, raising final-task success from approximately 0% to 90-100% in a controlled sequential-decision benchmark. However, capability does not imply investigation intelligence: an agent may have access to evidence without knowing when acquiring it is worth its cost. The original LLM-driven investigation trigger shows no significant advantage over simple heuristic policies in 10 of 12 paired comparisons; the two nominally significant differences do not survive Bonferroni correction. We introduce a validated agent-information oracle evaluating policies on only agent-available information, alongside a clairvoyant oracle giving an achievable ceiling. Using these oracles, we decompose policy regret into missed and unnecessary investigation and downstream action errors, revealing a systematic mismatch between investigation behavior and the oracle's value of information in a noisy, budget-scarce regime. Factorial experiments identify accumulated dismissed-failure count, rather than recency, as the primary behavioral factor associated with under-investigation in this setting. A targeted intervention further demonstrates that this behavior is manipulable, improving success from 91.7% to 94.0% in one regime (p = 0.020, n = 25), but not uniformly across budgets and noise levels. Because an explicit uncertainty scalar alone did not correct the remaining gap, we construct matched counterfactual histories from legitimate investigation-mediated evidence and localize, in the Hard, budget-2 regime, a narrow, repeatedly-observed belief-conditioned mismatch: behavior is consistent with an effective update threshold that is offset above the oracle's decision boundary. We further extend this diagnostic apparatus to an environment in which evidence must be attributed to the correct one of several tools, and find that simple heuristic baselines exhibit cross-tool evidence contamination that is absent in the tested LLM agents under the conditions evaluated - a result that is robust across models, a hard lexical confound, a vocabulary-domain transfer, and randomized phrasing, subject to an explicit, model-dependent boundary condition under simultaneous multi-source distractor pressure. These results suggest the central difficulty is not access to evidence or representation of uncertainty, but the selective, budget-aware conversion of uncertainty into appropriate investigation and update decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.