Bayesian Evidence Fusion for Uncertainty Quantification of AI Agents
Abstract
AI agents complete tasks through sequences of diverse actions: large language model (LLM) generations, tool calls, and revisions. Existing uncertainty quantification (UQ) techniques typically assess the reliability of individual LLM generations, leaving open how to aggregate and propagate uncertainty across different types of actions within an agent trajectory. We formulate UQ for LLM agents as Bayesian belief tracking over final task correctness, integrating generation-level uncertainty with tool outcomes through Bayesian updates. Our framework requires no changes to the agentic harness or LLM, and can be applied post hoc without additional agent executions, thus introducing negligible latency. We conduct an extensive evaluation across five LLMs, eight agent harnesses, and 15 agentic environments spanning coding, retrieval, and interactive tasks. The basic Bayesian fusion achieves the highest average outcome-ranking scores on coding and retrieval, while a variant that separately models trajectory length, history, and final-step evidence leads on interactive tasks. Bayesian fusion generalizes well across models, harnesses, and datasets within the same domain. For larger environment and domain shifts, adaptation with little or no labeled target data improves performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.