acceptodds
Under review as a conference paper at ICLR 2027

Explaining and Calibrating Dynamically Orchestrated LLM MAS through Stepwise Uncertainty Decomposition

Abstract

Motivated by the need for trustworthiness and interpretability, reliable uncertainty decomposition and calibration techniques provide transparency into the black-box behavior of Large Language Models (LLMs). However, extending this explainability to LLM-based Multi-Agent Systems (MAS) remains challenging, as stochastic interactions among agents and system-level control introduce complex, nested uncertainty. To address this challenge, we propose Explanation and Calibration through Orchestration for Multi-Agent Systems (EcoMAS), a framework that reveals stepwise information contributions in dynamically orchestrated MAS through uncertainty decomposition and optimizes orchestration policies to improve calibration. We model system output uncertainty using second-order Tsallis entropy and derive an exact decomposition into stepwise information contributions from orchestration and agent execution. We develop a paired sampling algorithm that efficiently produces unbiased estimates of these contributions with theoretical finite-sample concentration bounds. We further derive an unbiased calibration gradient estimator from the semantic increments underlying information attribution, reusing the explanation samples for orchestration policy optimization. We evaluate EcoMAS on MMLU-Pro, MATH-500, and ChaosNLI to assess estimation accuracy, explanation faithfulness, and orchestration calibration. The results show that paired sampling accurately estimates information contributions with small sampling budgets. Case studies and intervention evaluations qualitatively illustrate how the attributions explain answer revision and reinforcement, while quantitative faithfulness evaluations demonstrate that EcoMAS reliably identifies steps that strongly influence system output uncertainty. Finally, comparisons of orchestration optimization methods show that EcoMAS reduces calibration error while improving task accuracy, highlighting the association between semantic calibration and MAS output quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.